Netherlands Open Data Scraper - data.overheid.nl API
Pricing
from $6.23 / 1,000 datasets
Netherlands Open Data Scraper - data.overheid.nl API
Scrape all 20,380 datasets from data.overheid.nl, the national open data portal of the Netherlands: full DCAT-AP-DONL metadata, distributions, 4,667 data requests, registers and live link checks.
Pricing
from $6.23 / 1,000 datasets
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
8 days ago
Last modified
Categories
Share
Netherlands Open Data Scraper - data.overheid.nl CKAN API
Scrape every one of the 20,380 datasets on data.overheid.nl, the national open data portal of the Netherlands, with all 86 metadata fields the CKAN API carries. Each row is a flat DCAT-AP-DONL record: the data owner as an OWMS authority, the EU themes in Dutch and English, licence, access rights, update frequency, contact point, legal foundation and every file with its format. No login, no API key, no rate limit. Export to CSV, JSON, Excel, or XML.
The portal's own search shows you ten results a page and hides most of the record behind them, and its government open data API answers only in Dutch URIs: a licence comes back as http://creativecommons.org/licenses/by/4.0/deed.nl and a theme as .../owms/terms/Natuur_en_milieu. This Actor pages the whole open data catalogue in one run, resolves every URI to a readable name, and adds four registers the CKAN API does not expose at all, including the 4,667 public data requests Dutch citizens have filed asking the government to open a dataset.
| Who uses it | What they scrape data.overheid.nl for |
|---|---|
| Data engineers | A complete inventory of Netherlands government datasets to pipe into a warehouse |
| Open data researchers | Licence, access-rights and update-frequency coverage across 187 public bodies |
| Civic technologists | Live download links by format, checked, before building on top of them |
| Journalists | The 4,667 data requests and the government's own answer to each one |
| Compliance and policy teams | The 229 EU High Value Datasets and the 34 statutory basisregistraties |
What it does
This Actor reads the CKAN open data API behind data.overheid.nl and the five community registers on the site itself, and returns each entry as a flat row tagged with rowType:
- 🇳🇱 Datasets: all 20,380, with 86 fields each: title, description, national identifier, data owner and publisher, harvest organization, source catalogue, OWMS themes in Dutch and English, keywords, licence with an open-licence flag, access rights, availability status, update frequency, contact point with email, phone and website, issue and modification dates, temporal coverage, spatial extent and the legal foundation with its wetten.overheid.nl link.
- 📦 Files: around 74,600 distributions, 3.66 per dataset, with access and download URLs, the EU file-type URI resolved to a plain name, a machine-readable flag, media type, preview URL and licence.
- 📣 Data requests: all 4,667 public requests to open a dataset, with the phase they reached in Dutch and English, the body that was asked, and optionally the full request and the government's written answer.
- 🏛️ Registers: 1,671 public bodies with their governance layer (Rijk, Gemeente, Provincie, Waterschap), 91 curated groups, 36 source catalogues and the registered applications built on the data.
- 📚 Directories: 187 data owners, 99 themes, 40 formats, 8 licences, 23 source catalogues and 10,403 keywords, each with the dataset count and a working link to its slice of the portal.
- ✅ Flags the portal buries: 115 datasets flagged high value and 229 carrying an EU High Value Dataset category, 34 basisregistraties, 6,909 nationally-scoped datasets, 51 reference-data sets, and the 2,020 datasets catalogued but not openly downloadable.
- 🔎 Optional checks: fetch each file to see whether it still downloads (9.0 percent are not), and read its first bytes to see whether it really is what it claims (10.0 percent are not: 24 files declared JSON in a 469 file sample served an HTML page instead).
Results export to CSV, JSON, Excel, or XML, or stream from the API.
What you can do with data.overheid.nl data
🗂️ Build a complete inventory of Netherlands government data. One run with no filters returns all 20,380 datasets with their owners, licences, formats and update frequencies. Tick the file rows and you also get every download URL in the country's open data catalogue, ready to schedule.
🔗 Find the dead links before your users do. data.overheid.nl harvests 23 separate catalogues and records a link status on only 8 percent of files, so for the other 92 percent nobody has ever checked. Turn on the link check and each file comes back with its real HTTP status, redirect target, content type and size. In a 647 file sample, 9.0 percent were dead: mostly WFS and WMS services answering 500, plus spreadsheets behind a 403.
⚖️ Audit open data policy compliance. Filter to the 229 datasets in an EU High Value Dataset category, or to the 2,020 datasets whose access rights are not public, or to the 1,376 published under a closed licence, and you have the evidence for a coverage report in one export.
📣 Read what the public is asking for. The data request register is 4,667 requests from citizens, journalists and companies asking a named public body to open a dataset, each with the phase it reached and the answer it got. 336 were refused as not public and 1,970 ended with the government pointing at data that already existed.
Why choose this scraper
| What you get | |
|---|---|
| Full coverage | All 20,380 datasets, 86 fields each, plus five registers the CKAN API does not expose |
| Readable, not URIs | Every OWMS and EU URI resolved to a name, themes in Dutch and English |
| Filters measured against the live totals | 28 filters, each one checked against the catalogue count, with the real number in the dropdown |
| Clean rows | ISO dates, real numbers, Yes/No booleans, "Not Disclosed" where the portal withheld a value, zero always-empty columns |
| A register nobody else carries | 4,667 public data requests with the government's own answer |
| Priced per row | You pay for rows written, and the optional checks only when they returned something |
How it compares
One other Actor covers this exact source, and two more cover CKAN portals generically. The generic ones are cheaper per row and give you what a bare CKAN call gives you: an id, a name, a few tags. The named data.overheid.nl Actor lists six fields in its own store description. This Actor costs more per row than either, and the difference buys the other 80 fields, the URI-to-name resolution the portal makes you do yourself, the five community registers, and filters that were each measured against the live catalogue instead of copied from CKAN documentation.
| This Actor | benthepythondev/netherlands-data-overheid-scraper | Generic CKAN Actors | The portal's own API | |
|---|---|---|---|---|
| Fields per dataset | 86 | 6, per its own description | Whatever CKAN returns raw | Raw URIs only |
| Filters | 28, measured against totals | None documented | Free text | You write the Solr query |
| Data requests | 4,667 | No | No | Not exposed |
| Registers | 5 | No | No | Only organizations |
| Link checking | Yes, optional | No | No | No |
| Users in the last 30 days | New | 1 | 1 each | n/a |
What a Netherlands open dataset looks like
One real row from the verified run, unedited apart from truncating the description:
{"rowType": "dataset","url": "https://data.overheid.nl/dataset/wpozittenblijvers-v1","apiUrl": "https://data.overheid.nl/data/api/3/action/package_show?id=wpozittenblijvers-v1","id": "7cc95d10-bb52-4211-bbe8-8a18bb6e0f0d","slug": "wpozittenblijvers-v1","title": "Zittenblijvers in het bo en sbo per vestiging","description": "Dit document beschrijft de bestanden over leerlingen in het Primair Onderwijs die via een API (A ...","identifier": "https://onderwijsdata.duo.nl/datasets/wpozittenblijvers-v1","alternateIdentifiers": [],"organization": "dienst-uitvoering-onderwijs","organizationTitle": "DUO","organizationId": "3b966f6a-4bd6-4f63-95b4-76adcdf31657","dataOwner": "http://standaarden.overheid.nl/owms/terms/Dienst_Uitvoering_Onderwijs","dataOwnerName": "Dienst Uitvoering Onderwijs","publisher": "http://standaarden.overheid.nl/owms/terms/Dienst_Uitvoering_Onderwijs","publisherName": "Dienst Uitvoering Onderwijs","sourceCatalog": "https://onderwijsdata.duo.nl","sourceCatalogName": "Dienst Uitvoering Onderwijs (DUO)","themes": ["http://standaarden.overheid.nl/owms/terms/Onderwijs_en_wetenschap"],"themeLabels": ["Onderwijs en wetenschap"],"themeLabelsEn": ["Education and science"],"keywords": ["Leerlingen"],"keywordCount": 1,"licence": "http://creativecommons.org/licenses/by/4.0/deed.nl","licenceLabel": "CC-BY (4.0)","licenceUrl": "http://creativecommons.org/licenses/by/4.0/deed.nl","isOpenLicence": "Yes","accessRights": "http://publications.europa.eu/resource/authority/access-right/PUBLIC","accessRightsLabel": "Public","datasetStatus": "http://data.overheid.nl/status/beschikbaar","datasetStatusLabel": "Available","restrictionsStatement": "Not Disclosed","updateFrequency": "http://publications.europa.eu/resource/authority/frequency/ANNUAL","updateFrequencyLabel": "Annual","languages": ["http://publications.europa.eu/resource/authority/language/NLD"],"languageLabels": ["Dutch"],"metadataLanguage": "http://publications.europa.eu/resource/authority/language/NLD","highValueDataset": "No","hvdCategories": [],"hvdCategoryLabels": [],"basisRegister": "No","nationalCoverage": "No","referenceData": "No","datasetQuality": "N/A","contactName": "Informatieproducten","contactTitle": "Not Disclosed","contactEmail": "informatieproducten@duo.nl","contactPhone": "Not Disclosed","contactWebsite": "Not Disclosed","contactAddress": "Not Disclosed","author": "Informatieproducten","authorEmail": "informatieproducten@duo.nl","maintainer": "Not Disclosed","maintainerEmail": "Not Disclosed","version": "1.0.0","versionNotes": "Not Disclosed","metadataCreated": "2020-04-02T22:20:24.932693","metadataModified": "2026-08-27T08:10:20.178083","modified": "2026-01-09T07:43:36","issued": "Not Disclosed","datePlanned": "Not Disclosed","temporalStart": "Not Disclosed","temporalEnd": "Not Disclosed","temporalLabel": "Not Disclosed","spatialValues": [],"spatialSchemes": [],"legalFoundationLabel": "Not Disclosed","legalFoundationRef": "Not Disclosed","legalFoundationUrl": "Not Disclosed","provenance": [],"documentation": [],"samples": [],"sources": [],"conformsTo": [],"relatedResources": [],"landingPage": "https://onderwijsdata.duo.nl/datasets/wpozittenblijvers-v1","syncChecksum": "Not Disclosed","extraFields": [],"changeType": "updated","state": "active","isPrivate": "No","fileCount": 1,"fileFormats": ["http://publications.europa.eu/resource/authority/file-type/CSV"],"fileFormatLabels": ["CSV"],"files": [{"name": "Aantal zittenblijvers bo en sbo per schoolvestiging","url": "https://onderwijsdata.duo.nl/dataset/f0c79c94-6ffb-44bf-a3d5-9a05aef4ade4/resource/df7297a9-70a3-4a3a-b4a6-d289c7d16014/download/brin6_zittenblijvers.csv","format": "CSV","position": 0}],"scrapedAt": "2026-08-27T19:34:28.746Z"}
Configure the run
Leave everything empty and the Actor sweeps the whole open data catalogue, newest change first. Every dropdown carries the live dataset count next to each value, so you can see what a filter is worth before you run it. Filters combine with AND; several values inside one filter combine with OR.
Every dataset, newest change first, with its files:
{"maxItems": 20000,"sortBy": "metadata_modified desc","includeDistributions": true}
Environmental datasets published as CSV under an open licence, checked to see whether they still download:
{"themes": ["http://standaarden.overheid.nl/owms/terms/Natuur_en_milieu"],"formats": ["http://publications.europa.eu/resource/authority/file-type/CSV"],"licences": ["http://creativecommons.org/licenses/by/4.0/deed.nl"],"includeDistributions": true,"includeLinkCheck": true,"maxResourcesPerDataset": 5,"maxItems": 2000}
The data request register with the government's full answer to each request:
{"datasetIds": [],"includeDataRequests": true,"includeDataRequestDetails": true,"dataRequestPhases": ["Data is niet openbaar"],"maxItems": 400}
Pricing
Pay per event. You are charged for rows the Actor actually wrote, plus a small page fee for paging the catalogue, and the optional checks only when they returned something.
| Event | Price |
|---|---|
| Actor start | $0.054, once per run |
| Catalogue page scanned | $0.004 per page of up to 500 datasets |
| Dataset | $0.007 |
| File | $0.003 |
| Register page scanned | $0.004 per page of 10 entries |
| Data request | $0.006 |
| Data request detail read | $0.006 |
| Application, catalogue, group, organization | $0.005 to $0.006 |
| Data owner, theme, format, licence, source catalogue | $0.003 |
| Keyword | $0.002 |
| File link checked | $0.004 |
| File bytes probed | $0.012 |
| Data owner profile attached | $0.006 |
| Datasets | What it costs |
|---|---|
| 100 | $0.76 |
| 1,000 | $7.06 |
| 10,000 | $70.13 |
Paid Apify plans get a discount on every event: 3.8 percent on Bronze, 7.4 percent on Silver, 11 percent on Gold and above.
Free users
Free Apify accounts get 10 rows per run as a preview, which is enough to see every column and decide. Any paid plan lifts the cap to whatever you set in Max Items, up to 1,000,000. Upgrade here.
Run it
- Create a free Apify account. New accounts come with $5 in free credit, which is about 700 datasets.
- Open the Actor, leave the input empty for a full sweep or pick a theme, a data owner or a format from the dropdowns.
- Click Start. A 10-row preview finishes in about 2 seconds; a full 20,380-dataset sweep with files runs at roughly 267 rows per second.
- Download the results as CSV, JSON, Excel or XML, or pull them from the dataset API.
Use with AI agents (MCP)
claude mcp add apify --transport http https://mcp.apify.com --header "Authorization: Bearer YOUR_APIFY_TOKEN"
Then ask your agent in plain language:
- "List every Netherlands open dataset about air quality that is published as CSV under an open licence."
- "Which Dutch public bodies publish the most datasets, and how many of theirs are not openly downloadable?"
- "Show me the data requests where the government refused because the data is not public."
Troubleshooting
No results at all. A filter value the portal does not know returns zero rows rather than an error. Every dropdown here only offers values that exist, so the usual cause is combining two filters that have no overlap, for example a theme and a data owner that never publish together. Clear one filter and run again.
Fewer rows than I asked for. Max Items is a single budget shared by every row type. If you tick the extra registers, the datasets are collected first and can use the whole budget before the registers start. Raise Max Items or run a register on its own with no dataset filters.
A field is empty on every row. Most DCAT fields are optional upstream and the portal leaves them blank. legalFoundationLabel is filled on 3 percent of datasets, datasetQuality on 9 percent, spatialValues on under 1 percent. "Not Disclosed" means the portal withheld it; "N/A" means it does not apply. No column is empty on every dataset in the catalogue.
No community column on my rows. The Communities filter works but does not export: the portal indexes the community list for search and never returns it on the dataset record.
The run is slow. Plain dataset collection runs at about 267 rows per second, so 8,000 rows take 28 seconds. The link check and the byte probe fetch files from 23 different government servers, some of which take seconds to answer: 1,500 rows with both on and three files checked per dataset took 560 seconds. Lower "Files to check per dataset" to speed it up.
FAQ
| Question | Answer |
|---|---|
| Do I need an API key for data.overheid.nl? | No. The open data API is anonymous, and so is every page this Actor reads. |
| Is this the official CKAN API? | It reads the portal's CKAN API and its own website. It is not run by or affiliated with data.overheid.nl. |
| How many datasets are there? | 20,380 as of 27 August 2026, across 26 harvest organizations and 23 source catalogues. |
| Is the data in Dutch? | The records are, since the portal is Dutch-first: 18,956 are in Dutch, 1,389 in English, 35 in German. Themes, statuses, access rights, licences, formats and data request phases are also given in English. |
| Can I get only the machine-readable files? | Yes. Filter by format, and every file row also carries a machineReadable flag. |
| What is a basisregistratie? | One of the ten statutory Dutch national registers. 34 datasets are flagged as belonging to one, and there is a filter for them. |
| What are EU High Value Datasets? | Categories from the EU Open Data Directive that member states must publish for free. 229 datasets carry one, in five categories. |
| Does it check that the downloads work? | Optionally. The portal itself records a link status on only 8 percent of files; the link check does the other 92 percent, and found 9.0 percent of a 647 file sample dead. OGC services are asked for their GetCapabilities, so a live WFS is not reported as broken. |
| How fresh is it? | Live. Every run reads the portal at that moment, and 5,825 catalogue records changed in the last 30 days. |
| Can I scrape one dataset by URL? | Yes. Paste data.overheid.nl dataset or community URLs into Dataset or community URLs, or put CKAN slugs into Dataset IDs. |
Related actors
- Germany Open Data Scraper - GovData.de
- Italy Open Data Scraper - dati.gov.it
- CBS Netherlands Statistics Scraper
- EU Funding and Tenders Scraper
- TED Europa Tenders Scraper
Browse the full ParseForge collection for more scrapers.
Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.
Disclaimer: this is an unofficial tool, not affiliated with or endorsed by data.overheid.nl, KOOP or the government of the Netherlands. It collects only data that is already published openly, without logging in and without circumventing any access control. Records may contain names and contact details of public officials acting in their professional capacity; if you process them, GDPR, CCPA and PIPL obligations are yours as the data controller.
