# Netherlands Open Data Scraper - data.overheid.nl API (`parseforge/data-overheid-netherlands-open-data-scraper`) Actor

Scrape all 20,380 datasets from data.overheid.nl, the national open data portal of the Netherlands: full DCAT-AP-DONL metadata, distributions, 4,667 data requests, registers and live link checks.

- **URL**: https://apify.com/parseforge/data-overheid-netherlands-open-data-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Business
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $6.23 / 1,000 datasets

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner.jpg)](https://apify.com/parseforge?fpr=vmoqkp)

### Netherlands Open Data Scraper - data.overheid.nl CKAN API

**Scrape every one of the 20,380 datasets on data.overheid.nl, the national open data portal of the Netherlands, with all 86 metadata fields the CKAN API carries.** Each row is a flat DCAT-AP-DONL record: the data owner as an OWMS authority, the EU themes in Dutch and English, licence, access rights, update frequency, contact point, legal foundation and every file with its format. No login, no API key, no rate limit. Export to CSV, JSON, Excel, or XML.

The portal's own search shows you ten results a page and hides most of the record behind them, and its government open data API answers only in Dutch URIs: a licence comes back as `http://creativecommons.org/licenses/by/4.0/deed.nl` and a theme as `.../owms/terms/Natuur_en_milieu`. This Actor pages the whole open data catalogue in one run, resolves every URI to a readable name, and adds four registers the CKAN API does not expose at all, including the 4,667 public data requests Dutch citizens have filed asking the government to open a dataset.

| Who uses it | What they scrape data.overheid.nl for |
|---|---|
| Data engineers | A complete inventory of Netherlands government datasets to pipe into a warehouse |
| Open data researchers | Licence, access-rights and update-frequency coverage across 187 public bodies |
| Civic technologists | Live download links by format, checked, before building on top of them |
| Journalists | The 4,667 data requests and the government's own answer to each one |
| Compliance and policy teams | The 229 EU High Value Datasets and the 34 statutory basisregistraties |

### What it does

This Actor reads the CKAN open data API behind data.overheid.nl and the five community registers on the site itself, and returns each entry as a flat row tagged with `rowType`:

- 🇳🇱 **Datasets:** all 20,380, with 86 fields each: title, description, national identifier, data owner and publisher, harvest organization, source catalogue, OWMS themes in Dutch and English, keywords, licence with an open-licence flag, access rights, availability status, update frequency, contact point with email, phone and website, issue and modification dates, temporal coverage, spatial extent and the legal foundation with its wetten.overheid.nl link.
- 📦 **Files:** around 74,600 distributions, 3.66 per dataset, with access and download URLs, the EU file-type URI resolved to a plain name, a machine-readable flag, media type, preview URL and licence.
- 📣 **Data requests:** all 4,667 public requests to open a dataset, with the phase they reached in Dutch and English, the body that was asked, and optionally the full request and the government's written answer.
- 🏛️ **Registers:** 1,671 public bodies with their governance layer (Rijk, Gemeente, Provincie, Waterschap), 91 curated groups, 36 source catalogues and the registered applications built on the data.
- 📚 **Directories:** 187 data owners, 99 themes, 40 formats, 8 licences, 23 source catalogues and 10,403 keywords, each with the dataset count and a working link to its slice of the portal.
- ✅ **Flags the portal buries:** 115 datasets flagged high value and 229 carrying an EU High Value Dataset category, 34 basisregistraties, 6,909 nationally-scoped datasets, 51 reference-data sets, and the 2,020 datasets catalogued but not openly downloadable.
- 🔎 **Optional checks:** fetch each file to see whether it still downloads (9.0 percent are not), and read its first bytes to see whether it really is what it claims (10.0 percent are not: 24 files declared JSON in a 469 file sample served an HTML page instead).

Results export to CSV, JSON, Excel, or XML, or stream from the API.

### What you can do with data.overheid.nl data

**🗂️ Build a complete inventory of Netherlands government data.** One run with no filters returns all 20,380 datasets with their owners, licences, formats and update frequencies. Tick the file rows and you also get every download URL in the country's open data catalogue, ready to schedule.

**🔗 Find the dead links before your users do.** data.overheid.nl harvests 23 separate catalogues and records a link status on only 8 percent of files, so for the other 92 percent nobody has ever checked. Turn on the link check and each file comes back with its real HTTP status, redirect target, content type and size. In a 647 file sample, 9.0 percent were dead: mostly WFS and WMS services answering 500, plus spreadsheets behind a 403.

**⚖️ Audit open data policy compliance.** Filter to the 229 datasets in an EU High Value Dataset category, or to the 2,020 datasets whose access rights are not public, or to the 1,376 published under a closed licence, and you have the evidence for a coverage report in one export.

**📣 Read what the public is asking for.** The data request register is 4,667 requests from citizens, journalists and companies asking a named public body to open a dataset, each with the phase it reached and the answer it got. 336 were refused as not public and 1,970 ended with the government pointing at data that already existed.

### Why choose this scraper

| | What you get |
|---|---|
| Full coverage | All 20,380 datasets, 86 fields each, plus five registers the CKAN API does not expose |
| Readable, not URIs | Every OWMS and EU URI resolved to a name, themes in Dutch and English |
| Filters measured against the live totals | 28 filters, each one checked against the catalogue count, with the real number in the dropdown |
| Clean rows | ISO dates, real numbers, Yes/No booleans, `"Not Disclosed"` where the portal withheld a value, zero always-empty columns |
| A register nobody else carries | 4,667 public data requests with the government's own answer |
| Priced per row | You pay for rows written, and the optional checks only when they returned something |

### How it compares

One other Actor covers this exact source, and two more cover CKAN portals generically. The generic ones are cheaper per row and give you what a bare CKAN call gives you: an id, a name, a few tags. The named data.overheid.nl Actor lists six fields in its own store description. This Actor costs more per row than either, and the difference buys the other 80 fields, the URI-to-name resolution the portal makes you do yourself, the five community registers, and filters that were each measured against the live catalogue instead of copied from CKAN documentation.

| | This Actor | benthepythondev/netherlands-data-overheid-scraper | Generic CKAN Actors | The portal's own API |
|---|---|---|---|---|
| Fields per dataset | 86 | 6, per its own description | Whatever CKAN returns raw | Raw URIs only |
| Filters | 28, measured against totals | None documented | Free text | You write the Solr query |
| Data requests | 4,667 | No | No | Not exposed |
| Registers | 5 | No | No | Only organizations |
| Link checking | Yes, optional | No | No | No |
| Users in the last 30 days | New | 1 | 1 each | n/a |

### What a Netherlands open dataset looks like

One real row from the verified run, unedited apart from truncating the description:

```json
{
  "rowType": "dataset",
  "url": "https://data.overheid.nl/dataset/wpozittenblijvers-v1",
  "apiUrl": "https://data.overheid.nl/data/api/3/action/package_show?id=wpozittenblijvers-v1",
  "id": "7cc95d10-bb52-4211-bbe8-8a18bb6e0f0d",
  "slug": "wpozittenblijvers-v1",
  "title": "Zittenblijvers in het bo en sbo per vestiging",
  "description": "Dit document beschrijft de bestanden over leerlingen in het Primair Onderwijs die via een API (A ...",
  "identifier": "https://onderwijsdata.duo.nl/datasets/wpozittenblijvers-v1",
  "alternateIdentifiers": [],
  "organization": "dienst-uitvoering-onderwijs",
  "organizationTitle": "DUO",
  "organizationId": "3b966f6a-4bd6-4f63-95b4-76adcdf31657",
  "dataOwner": "http://standaarden.overheid.nl/owms/terms/Dienst_Uitvoering_Onderwijs",
  "dataOwnerName": "Dienst Uitvoering Onderwijs",
  "publisher": "http://standaarden.overheid.nl/owms/terms/Dienst_Uitvoering_Onderwijs",
  "publisherName": "Dienst Uitvoering Onderwijs",
  "sourceCatalog": "https://onderwijsdata.duo.nl",
  "sourceCatalogName": "Dienst Uitvoering Onderwijs (DUO)",
  "themes": [
    "http://standaarden.overheid.nl/owms/terms/Onderwijs_en_wetenschap"
  ],
  "themeLabels": [
    "Onderwijs en wetenschap"
  ],
  "themeLabelsEn": [
    "Education and science"
  ],
  "keywords": [
    "Leerlingen"
  ],
  "keywordCount": 1,
  "licence": "http://creativecommons.org/licenses/by/4.0/deed.nl",
  "licenceLabel": "CC-BY (4.0)",
  "licenceUrl": "http://creativecommons.org/licenses/by/4.0/deed.nl",
  "isOpenLicence": "Yes",
  "accessRights": "http://publications.europa.eu/resource/authority/access-right/PUBLIC",
  "accessRightsLabel": "Public",
  "datasetStatus": "http://data.overheid.nl/status/beschikbaar",
  "datasetStatusLabel": "Available",
  "restrictionsStatement": "Not Disclosed",
  "updateFrequency": "http://publications.europa.eu/resource/authority/frequency/ANNUAL",
  "updateFrequencyLabel": "Annual",
  "languages": [
    "http://publications.europa.eu/resource/authority/language/NLD"
  ],
  "languageLabels": [
    "Dutch"
  ],
  "metadataLanguage": "http://publications.europa.eu/resource/authority/language/NLD",
  "highValueDataset": "No",
  "hvdCategories": [],
  "hvdCategoryLabels": [],
  "basisRegister": "No",
  "nationalCoverage": "No",
  "referenceData": "No",
  "datasetQuality": "N/A",
  "contactName": "Informatieproducten",
  "contactTitle": "Not Disclosed",
  "contactEmail": "informatieproducten@duo.nl",
  "contactPhone": "Not Disclosed",
  "contactWebsite": "Not Disclosed",
  "contactAddress": "Not Disclosed",
  "author": "Informatieproducten",
  "authorEmail": "informatieproducten@duo.nl",
  "maintainer": "Not Disclosed",
  "maintainerEmail": "Not Disclosed",
  "version": "1.0.0",
  "versionNotes": "Not Disclosed",
  "metadataCreated": "2020-04-02T22:20:24.932693",
  "metadataModified": "2026-08-27T08:10:20.178083",
  "modified": "2026-01-09T07:43:36",
  "issued": "Not Disclosed",
  "datePlanned": "Not Disclosed",
  "temporalStart": "Not Disclosed",
  "temporalEnd": "Not Disclosed",
  "temporalLabel": "Not Disclosed",
  "spatialValues": [],
  "spatialSchemes": [],
  "legalFoundationLabel": "Not Disclosed",
  "legalFoundationRef": "Not Disclosed",
  "legalFoundationUrl": "Not Disclosed",
  "provenance": [],
  "documentation": [],
  "samples": [],
  "sources": [],
  "conformsTo": [],
  "relatedResources": [],
  "landingPage": "https://onderwijsdata.duo.nl/datasets/wpozittenblijvers-v1",
  "syncChecksum": "Not Disclosed",
  "extraFields": [],
  "changeType": "updated",
  "state": "active",
  "isPrivate": "No",
  "fileCount": 1,
  "fileFormats": [
    "http://publications.europa.eu/resource/authority/file-type/CSV"
  ],
  "fileFormatLabels": [
    "CSV"
  ],
  "files": [
    {
      "name": "Aantal zittenblijvers bo en sbo per schoolvestiging",
      "url": "https://onderwijsdata.duo.nl/dataset/f0c79c94-6ffb-44bf-a3d5-9a05aef4ade4/resource/df7297a9-70a3-4a3a-b4a6-d289c7d16014/download/brin6_zittenblijvers.csv",
      "format": "CSV",
      "position": 0
    }
  ],
  "scrapedAt": "2026-08-27T19:34:28.746Z"
}
```

### Configure the run

Leave everything empty and the Actor sweeps the whole open data catalogue, newest change first. Every dropdown carries the live dataset count next to each value, so you can see what a filter is worth before you run it. Filters combine with AND; several values inside one filter combine with OR.

Every dataset, newest change first, with its files:

```json
{
  "maxItems": 20000,
  "sortBy": "metadata_modified desc",
  "includeDistributions": true
}
```

Environmental datasets published as CSV under an open licence, checked to see whether they still download:

```json
{
  "themes": ["http://standaarden.overheid.nl/owms/terms/Natuur_en_milieu"],
  "formats": ["http://publications.europa.eu/resource/authority/file-type/CSV"],
  "licences": ["http://creativecommons.org/licenses/by/4.0/deed.nl"],
  "includeDistributions": true,
  "includeLinkCheck": true,
  "maxResourcesPerDataset": 5,
  "maxItems": 2000
}
```

The data request register with the government's full answer to each request:

```json
{
  "datasetIds": [],
  "includeDataRequests": true,
  "includeDataRequestDetails": true,
  "dataRequestPhases": ["Data is niet openbaar"],
  "maxItems": 400
}
```

### Pricing

Pay per event. You are charged for rows the Actor actually wrote, plus a small page fee for paging the catalogue, and the optional checks only when they returned something.

| Event | Price |
|---|---|
| Actor start | $0.054, once per run |
| Catalogue page scanned | $0.004 per page of up to 500 datasets |
| Dataset | $0.007 |
| File | $0.003 |
| Register page scanned | $0.004 per page of 10 entries |
| Data request | $0.006 |
| Data request detail read | $0.006 |
| Application, catalogue, group, organization | $0.005 to $0.006 |
| Data owner, theme, format, licence, source catalogue | $0.003 |
| Keyword | $0.002 |
| File link checked | $0.004 |
| File bytes probed | $0.012 |
| Data owner profile attached | $0.006 |

| Datasets | What it costs |
|---|---|
| 100 | $0.76 |
| 1,000 | $7.06 |
| 10,000 | $70.13 |

Paid Apify plans get a discount on every event: 3.8 percent on Bronze, 7.4 percent on Silver, 11 percent on Gold and above.

### Free users

Free Apify accounts get 10 rows per run as a preview, which is enough to see every column and decide. Any paid plan lifts the cap to whatever you set in Max Items, up to 1,000,000. [Upgrade here](https://apify.com/pricing?fpr=vmoqkp).

### Run it

1. Create a free Apify account. New accounts come with $5 in free credit, which is about 700 datasets.
2. Open the Actor, leave the input empty for a full sweep or pick a theme, a data owner or a format from the dropdowns.
3. Click Start. A 10-row preview finishes in about 2 seconds; a full 20,380-dataset sweep with files runs at roughly 267 rows per second.
4. Download the results as CSV, JSON, Excel or XML, or pull them from the dataset API.

### Use with AI agents (MCP)

```
claude mcp add apify --transport http https://mcp.apify.com --header "Authorization: Bearer YOUR_APIFY_TOKEN"
```

Then ask your agent in plain language:

- "List every Netherlands open dataset about air quality that is published as CSV under an open licence."
- "Which Dutch public bodies publish the most datasets, and how many of theirs are not openly downloadable?"
- "Show me the data requests where the government refused because the data is not public."

### Troubleshooting

**No results at all.** A filter value the portal does not know returns zero rows rather than an error. Every dropdown here only offers values that exist, so the usual cause is combining two filters that have no overlap, for example a theme and a data owner that never publish together. Clear one filter and run again.

**Fewer rows than I asked for.** Max Items is a single budget shared by every row type. If you tick the extra registers, the datasets are collected first and can use the whole budget before the registers start. Raise Max Items or run a register on its own with no dataset filters.

**A field is empty on every row.** Most DCAT fields are optional upstream and the portal leaves them blank. `legalFoundationLabel` is filled on 3 percent of datasets, `datasetQuality` on 9 percent, `spatialValues` on under 1 percent. `"Not Disclosed"` means the portal withheld it; `"N/A"` means it does not apply. No column is empty on every dataset in the catalogue.

**No community column on my rows.** The Communities filter works but does not export: the portal indexes the community list for search and never returns it on the dataset record.

**The run is slow.** Plain dataset collection runs at about 267 rows per second, so 8,000 rows take 28 seconds. The link check and the byte probe fetch files from 23 different government servers, some of which take seconds to answer: 1,500 rows with both on and three files checked per dataset took 560 seconds. Lower "Files to check per dataset" to speed it up.

### FAQ

| Question | Answer |
|---|---|
| Do I need an API key for data.overheid.nl? | No. The open data API is anonymous, and so is every page this Actor reads. |
| Is this the official CKAN API? | It reads the portal's CKAN API and its own website. It is not run by or affiliated with data.overheid.nl. |
| How many datasets are there? | 20,380 as of 27 August 2026, across 26 harvest organizations and 23 source catalogues. |
| Is the data in Dutch? | The records are, since the portal is Dutch-first: 18,956 are in Dutch, 1,389 in English, 35 in German. Themes, statuses, access rights, licences, formats and data request phases are also given in English. |
| Can I get only the machine-readable files? | Yes. Filter by format, and every file row also carries a `machineReadable` flag. |
| What is a basisregistratie? | One of the ten statutory Dutch national registers. 34 datasets are flagged as belonging to one, and there is a filter for them. |
| What are EU High Value Datasets? | Categories from the EU Open Data Directive that member states must publish for free. 229 datasets carry one, in five categories. |
| Does it check that the downloads work? | Optionally. The portal itself records a link status on only 8 percent of files; the link check does the other 92 percent, and found 9.0 percent of a 647 file sample dead. OGC services are asked for their GetCapabilities, so a live WFS is not reported as broken. |
| How fresh is it? | Live. Every run reads the portal at that moment, and 5,825 catalogue records changed in the last 30 days. |
| Can I scrape one dataset by URL? | Yes. Paste data.overheid.nl dataset or community URLs into Dataset or community URLs, or put CKAN slugs into Dataset IDs. |

### Related actors

- [Germany Open Data Scraper - GovData.de](https://apify.com/parseforge/govdata-germany-scraper?fpr=vmoqkp)
- [Italy Open Data Scraper - dati.gov.it](https://apify.com/parseforge/dati-gov-it-italy-open-data-scraper?fpr=vmoqkp)
- [CBS Netherlands Statistics Scraper](https://apify.com/parseforge/cbs-netherlands-statistics-scraper?fpr=vmoqkp)
- [EU Funding and Tenders Scraper](https://apify.com/parseforge/eu-funding-tenders-scraper?fpr=vmoqkp)
- [TED Europa Tenders Scraper](https://apify.com/parseforge/ted-europa-tenders-scraper?fpr=vmoqkp)

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

Disclaimer: this is an unofficial tool, not affiliated with or endorsed by data.overheid.nl, KOOP or the government of the Netherlands. It collects only data that is already published openly, without logging in and without circumventing any access control. Records may contain names and contact details of public officials acting in their professional capacity; if you process them, GDPR, CCPA and PIPL obligations are yours as the data controller.

# Actor input Schema

## `searchTerms` (type: `array`):

Free-text queries run against title, description and keywords. Dutch works best: verkeer, bevolking, milieu, parkeren. Leave empty to sweep the whole catalogue of 20,380 datasets.

## `startUrls` (type: `array`):

Specific data.overheid.nl pages, for example https://data.overheid.nl/dataset/wpozittenblijvers-v1 or https://data.overheid.nl/community/datarequest/schengen-entry-ban. When you supply these, only these pages are scraped and the filters below are ignored.

## `datasetIds` (type: `array`):

CKAN dataset names or UUIDs, one per line, for example wpozittenblijvers-v1. Same effect as Dataset URLs, without the URL.

## `maxItems` (type: `integer`):

Free users: limited to 10 items (preview). Paid users: up to 1,000,000.

## `themes` (type: `array`):

Dutch government OWMS themes. Sub-themes are listed under their parent. The catalogue stores the leaf theme only, so picking a parent does NOT include its children: tick both if you want both.

## `dataOwners` (type: `array`):

The public body legally responsible for the data (OWMS authority). All 188 values the portal offers are listed with their dataset count; the catalogue index itself resolves 187 of them.

## `organizations` (type: `array`):

The portal account a dataset was loaded under, which is usually the source portal rather than the data owner. 26 exist.

## `sourceCatalogs` (type: `array`):

The regional, municipal or ministerial portal a dataset was harvested from. 23 feed the national catalogue.

## `formats` (type: `array`):

Keep only datasets that publish at least one file in these formats. The index stores the full EU file-type URI, so the short spelling CSV matches nothing and every value here is the URI the portal really uses.

## `licences` (type: `array`):

All 8 licences the portal records. CC-BY 4.0, CC-0 and Public domain together cover 18,349 of the 20,380 datasets.

## `statuses` (type: `array`):

Whether the data behind the record is actually available. 20,108 datasets are Available; the other three states are rare and worth isolating.

## `accessRights` (type: `array`):

DCAT access rights. 2,020 datasets are catalogued but not openly downloadable, which is exactly what a data-availability audit is looking for.

## `updateFrequencies` (type: `array`):

DCAT accrual periodicity. Populated on 30 percent of datasets; the rest declare nothing and are excluded by any choice here.

## `languages` (type: `array`):

Language the metadata is written in. The portal is Dutch-first: 18,956 records are Dutch, 1,389 English, 35 German.

## `hvdCategories` (type: `array`):

Categories from the EU Open Data Directive. Only 229 datasets carry one, and they are the ones the Directive obliges the Netherlands to publish for free.

## `communities` (type: `array`):

The four thematic communities data.overheid.nl runs. This one filters but does not export: the portal indexes the community list for search and never returns it on the dataset record, so there is no community column on the output rows. Counts here are the ones the index really matches, which are lower than the site shows.

## `keywords` (type: `array`):

Keyword tags exactly as the portal stores them, one per line, for example verkeer, natuur, bodem or water. 10,403 exist. Tick "Keywords directory" below to export the full list with counts.

## `minResources` (type: `integer`):

Keep only datasets that ship at least this many distributions. 6,754 datasets have 3 or more.

## `modifiedFrom` (type: `string`):

Only datasets whose catalogue record changed on or after this date (YYYY-MM-DD). 5,825 changed in the last 30 days.

## `modifiedTo` (type: `string`):

Only datasets whose catalogue record changed on or before this date (YYYY-MM-DD).

## `createdFrom` (type: `string`):

Only datasets first indexed on data.overheid.nl on or after this date (YYYY-MM-DD). 8,531 were added since 2024.

## `createdTo` (type: `string`):

Only datasets first indexed on data.overheid.nl on or before this date (YYYY-MM-DD).

## `issuedFrom` (type: `string`):

The publisher's own issue date, not the harvest date (YYYY-MM-DD). Populated on 49 percent of datasets.

## `issuedTo` (type: `string`):

Upper bound for the publisher's own issue date (YYYY-MM-DD).

## `onlyHighValue` (type: `boolean`):

Keep only the 115 datasets flagged as high value under the EU Open Data Directive.

## `onlyBasisRegister` (type: `boolean`):

Keep only the 34 datasets that belong to a Dutch basisregistratie, the ten statutory national registers.

## `onlyNationalCoverage` (type: `boolean`):

Keep only the 6,909 datasets that cover the whole country rather than one municipality or province.

## `onlyReferenceData` (type: `boolean`):

Keep only the 51 datasets marked as reference data (code lists and authoritative vocabularies).

## `customFilterQuery` (type: `string`):

Raw CKAN fq clause ANDed with everything above, for example num\_resources:\[10 TO \*] AND -organization:nationaalgeoregister-nl. Wrap any OR group in parentheses: a bare top-level OR returns the whole catalogue upstream. A field name that does not exist returns zero rows rather than an error.

## `sortBy` (type: `string`):

Only these six orders actually work upstream. Any other value is silently ignored by the portal and falls back to relevance.

## `dataRequestSearch` (type: `string`):

Free text over the data request register, for example verkeer. The register ignores any other search parameter without saying so, and 160 of the 4,667 requests mention verkeer.

## `dataRequestPhases` (type: `array`):

Where the request ended up. "Data is beschikbaar" means the government pointed at data that already exists; "Data is niet openbaar" means it refused.

## `dataRequestAuthorityKinds` (type: `array`):

Which layer of government the request was aimed at.

## `includeDistributions` (type: `boolean`):

Emit one extra row per downloadable file or service, with its own URL, EU format URI, media type, licence, preview URL and the link status the portal itself last recorded. Datasets average 4.1 files each.

## `includeDataRequests` (type: `boolean`):

Emit the 4,667 public data requests: what was asked for, which body was asked, the phase it reached and the status. Nothing else on the Store carries this register.

## `includeApplications` (type: `boolean`):

Emit the 7 registered applications built on Netherlands open data, with their type, maker, live URL and contact.

## `includeCatalogs` (type: `boolean`):

Emit the 36 source catalogues the portal harvests, each with its own portal URL, description and dataset count.

## `includeGroups` (type: `boolean`):

Emit the 91 curated dataset groups, for example the ten basisregistraties or the municipal High Value Data list, with their descriptions.

## `includeOrganizations` (type: `boolean`):

Emit the 1,671 public bodies in the portal register with their governance layer (Rijk, Gemeente, Provincie, Waterschap), description, logo and dataset count. This is a large register: raise Max Items before ticking it.

## `includeDataOwners` (type: `boolean`):

Emit one row per data owner that appears in your filtered result set, with the number of datasets it holds.

## `includeThemes` (type: `boolean`):

Emit all 99 OWMS themes and sub-themes with their Dutch and English names, their parent theme and their dataset counts.

## `includeFormats` (type: `boolean`):

Emit every file format in the matched result set with its canonical name, machine-readable flag and dataset count.

## `includeLicences` (type: `boolean`):

Emit every licence in the matched result set with its plain-English label, whether it is open, and how many datasets declare it.

## `includeKeywords` (type: `boolean`):

Emit the keyword tags used by the datasets that match your filters, with counts, ranked. The catalogue holds 10,403 of them.

## `includeSourceCatalogs` (type: `boolean`):

Emit the source portals that feed the matched result set, with how many datasets each contributes.

## `includeDataRequestDetails` (type: `boolean`):

Open every data request and add the 15 fields the listing hides: the exact data asked for, the format and period wanted, the intended use and impact, what the requester already tried, the full government reply, the resolution and the body named as the possible data owner.

## `includeLinkCheck` (type: `boolean`):

Request each file and record the HTTP status, redirect target, content type, byte size and latency. data.overheid.nl harvests 23 catalogues and stores a link status on only 8 percent of files, so for the other 92 percent nobody has ever checked.

## `includeFileProbe` (type: `boolean`):

Download the first 64 KB of each file and report what the bytes really are, whether that matches the declared format, the CSV delimiter and the column headers.

## `includeAuthorityProfile` (type: `boolean`):

Add the owning body's governance layer (Rijk, Gemeente, Provincie, Waterschap), description, logo and total dataset count to every dataset row. Each owner is fetched once per run and reused.

## `maxResourcesPerDataset` (type: `integer`):

Upper bound on how many files the two checks above touch per dataset. Only used when one of them is on.

## Actor input object example

```json
{
  "searchTerms": [],
  "startUrls": [],
  "datasetIds": [],
  "maxItems": 10,
  "themes": [],
  "dataOwners": [],
  "organizations": [],
  "sourceCatalogs": [],
  "formats": [],
  "licences": [],
  "statuses": [],
  "accessRights": [],
  "updateFrequencies": [],
  "languages": [],
  "hvdCategories": [],
  "communities": [],
  "keywords": [],
  "onlyHighValue": false,
  "onlyBasisRegister": false,
  "onlyNationalCoverage": false,
  "onlyReferenceData": false,
  "sortBy": "metadata_modified desc",
  "dataRequestPhases": [],
  "dataRequestAuthorityKinds": [],
  "includeDistributions": false,
  "includeDataRequests": false,
  "includeApplications": false,
  "includeCatalogs": false,
  "includeGroups": false,
  "includeOrganizations": false,
  "includeDataOwners": false,
  "includeThemes": false,
  "includeFormats": false,
  "includeLicences": false,
  "includeKeywords": false,
  "includeSourceCatalogs": false,
  "includeDataRequestDetails": false,
  "includeLinkCheck": false,
  "includeFileProbe": false,
  "includeAuthorityProfile": false,
  "maxResourcesPerDataset": 5
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `csv` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [],
    "startUrls": [],
    "datasetIds": [],
    "maxItems": 10,
    "themes": [],
    "dataOwners": [],
    "organizations": [],
    "sourceCatalogs": [],
    "formats": [],
    "licences": [],
    "statuses": [],
    "accessRights": [],
    "updateFrequencies": [],
    "languages": [],
    "hvdCategories": [],
    "communities": [],
    "keywords": [],
    "sortBy": "metadata_modified desc",
    "dataRequestPhases": [],
    "dataRequestAuthorityKinds": [],
    "maxResourcesPerDataset": 5
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/data-overheid-netherlands-open-data-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerms": [],
    "startUrls": [],
    "datasetIds": [],
    "maxItems": 10,
    "themes": [],
    "dataOwners": [],
    "organizations": [],
    "sourceCatalogs": [],
    "formats": [],
    "licences": [],
    "statuses": [],
    "accessRights": [],
    "updateFrequencies": [],
    "languages": [],
    "hvdCategories": [],
    "communities": [],
    "keywords": [],
    "sortBy": "metadata_modified desc",
    "dataRequestPhases": [],
    "dataRequestAuthorityKinds": [],
    "maxResourcesPerDataset": 5,
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/data-overheid-netherlands-open-data-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [],
  "startUrls": [],
  "datasetIds": [],
  "maxItems": 10,
  "themes": [],
  "dataOwners": [],
  "organizations": [],
  "sourceCatalogs": [],
  "formats": [],
  "licences": [],
  "statuses": [],
  "accessRights": [],
  "updateFrequencies": [],
  "languages": [],
  "hvdCategories": [],
  "communities": [],
  "keywords": [],
  "sortBy": "metadata_modified desc",
  "dataRequestPhases": [],
  "dataRequestAuthorityKinds": [],
  "maxResourcesPerDataset": 5
}' |
apify call parseforge/data-overheid-netherlands-open-data-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/data-overheid-netherlands-open-data-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/lzskoItjyqodgXfd0/builds/8gidjXj2YLfP1kXCS/openapi.json
