Ireland Open Data Scraper - data.gov.ie Portal avatar

Ireland Open Data Scraper - data.gov.ie Portal

Pricing

from $6.23 / 1,000 datasets

Go to Apify Store
Ireland Open Data Scraper - data.gov.ie Portal

Ireland Open Data Scraper - data.gov.ie Portal

Scrape all 22,666 datasets from Ireland's open data portal data.gov.ie: full DCAT metadata, distributions, 174 publishers, EU high-value flags, openness scores and DataStore column schemas.

Pricing

from $6.23 / 1,000 datasets

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

ParseForge

Ireland Open Data Scraper - data.gov.ie Portal API

Every one of the 22,666 datasets on data.gov.ie, Ireland's open data portal, with 69 fields on each row. Each row carries the title, description, publishing body, DCAT theme, tags, licence, EU high-value flag, update frequency, dates, spatial and temporal coverage, contact point, harvest provenance and every downloadable file with its format and direct URL. No login, no API key, no rate limit to negotiate. Exports to CSV, JSON, Excel and XML.

data.gov.ie publishes a CKAN API, and you can talk to it yourself. What you cannot do from a single call is page 22,666 datasets without tripping its HTTP 409 on a second filter parameter, know that res_format:csv returns zero while res_format:CSV returns 20,237, or discover that the portal's 0 to 5 star openness rating only ever rides on the record view and never on the search results. This Actor already knows all of that, and adds four things the portal will not hand you in one place: the openness score, the real column schema and row count of every file loaded into the portal DataStore, a live check on whether each file still downloads, and a normalised publisher profile.

Who uses itWhat they scrape data.gov.ie for
Data engineersBuild a searchable mirror of Irish public data and refresh it from the change feed
Compliance and policy teamsTrack which bodies meet the EU high-value dataset obligation under Regulation 2023/138
Civic tech and journalistsFind the councils and agencies that publish, and the ones that publish and then let the links rot
AI and RAG buildersFeed 22,666 described, licensed, machine-readable Irish datasets into a retrieval index
ResearchersPull publisher, theme, licence and frequency directories to measure how Irish open data actually behaves

What it does

  • 📇 Full dataset record, 69 fields: id, slug, title, description in raw and plain-text form, publisher with its own description, theme, tags, licence with a plain-English label, dates, version, contact point, harvest collection and source.
  • 🇪🇺 EU high-value block: the high_value_dataset flag, the Regulation (EU) 2023/138 category resolved from its data.europa.eu URI, and the applicable legislation the publisher cites.
  • 🗺️ Geospatial block: spatial coverage name, bounding box computed from the stored geometry, coordinate reference systems, spatial resolution and vertical extent with its datum.
  • 📄 A row per file with format normalised across 132 raw spellings, MIME type, byte size, checksum, direct download URL and whether the portal loaded it into its DataStore.
  • Openness score, the portal's own 0 to 5 star rating with the reason it was given, present on every dataset and reachable nowhere in the search API.
  • 🧮 DataStore column schemas: for files the portal has loaded, the real column names, per-column SQL types, the true row count and a sample of rows.
  • 🔗 Live link check on each file URL: HTTP status, redirect target, content type and byte size, because the portal harvests 31 external feeds and never revalidates them.
  • 🏛️ Twelve directory row types: publishers, themes, EU high-value categories, licences, update frequencies, harvest collections, formats, tags, harvest sources, showcases, public data requests and the portal change feed.
  • 🎛️ Twenty-three filters measured against live totals, plus free-text search, a raw Solr escape hatch and thirteen verified sort orders.

What you can do with data.gov.ie data

Audit EU high-value compliance. Ireland flags 383 datasets as high value under Implementing Regulation (EU) 2023/138. Another 4,217 carry the Geospatial category URI from that same regulation while their high-value flag still says false. Filter to onlyHighValue, or to the Geospatial category, pull the publisher and the applicable legislation on each row, and you have the gap list in one run.

Find the rot before your pipeline does. Turn on the link check and every file URL is fetched from the publisher's own server. On 4,490 files checked across 1,531 datasets, 32 no longer downloaded. Those are ArcGIS Hubs and council portals the national portal copied years ago and never rechecked.

Know what is inside a file before you download it. 19.4% of files are loaded into the portal DataStore. For those, the DataStore probe returns the actual column names, their SQL types and the row count: 3,049 rows over 11 typed columns for the primary-school allocations CSV, without pulling a single byte of the file.

Watch the portal instead of polling it. The change feed row type returns recent dataset changes with timestamps. Run it on a schedule, diff against your last run, and refetch only what moved.

Why choose this scraper

What you get
CoverageAll 22,666 datasets, 174 publishers, 31 harvest sources, 14 showcases, 112 public data requests, and the top 2,000 of the portal's 22,742 tags
Fields69 on the base dataset row, 85 with all four enrichment blocks on, 40 on a file row
Filters23 filters plus free-text search, every one measured against the unfiltered total before it shipped
CorrectnessFormat spellings expanded case-correctly, invalid sorts rejected instead of silently ignored, filters ANDed into one clause so the portal never answers 409
EnrichmentOpenness score, DataStore column schemas, live link checks, publisher profiles
Cost controlFourteen row types and four enrichment blocks billed separately, so you pay for what you switch on

How it compares

Two other Actors cover data.gov.ie and one covers CKAN portals generically. All three are cheaper per row, at $2.00 to $2.50 per 1,000 against our $7.00, and all three had one user in the last 30 days. The difference is what a row is. benthepythondev/ireland-data-gov-packages-scraper takes one input, maxResults, and returns ids and names. benthepythondev/ireland-open-data-scraper takes two, a query and maxResults. straightforward_hydra/ckan-open-data-scraper is a generic CKAN reader that knows nothing Irish. None of them carry the EU high-value block, the openness score, the DataStore schemas or a link check, and none can filter by theme, publisher, licence, frequency, harvest collection or date range. If you want a list of dataset names, buy the cheap one. If you want the portal modelled, this is the one.

This Actorbenthepythondev/ireland-open-data-scraperstraightforward_hydra/ckan-open-data-scraper
Price per 1,000 rows$7.00$2.00$2.00
Input fields472generic CKAN
Filters measured against totals231 free-text querynone Ireland-specific
Row types1411
EU high-value blockYesNoNo
Openness scoreYesNoNo
DataStore column schemasYesNoNo
Live link checkYesNoNo

What a dataset looks like

One real, unedited row from the verified run, with all four enrichment blocks switched on:

{
"datasetId": "735d61de-6400-4b81-832d-8ff6cc84fa4f",
"name": "2026-2027-school-allocations",
"url": "https://data.gov.ie/dataset/2026-2027-school-allocations",
"title": "2026-2027 School Allocations",
"titleIrish": "Not Disclosed",
"description": "2026-2027 School Allocations. There are 3 files with allocations. Allocations are for Special Schools, Primary Schools and Post Primary Schools.",
"descriptionText": "2026-2027 School Allocations. There are 3 files with allocations. Allocations are for Special Schools, Primary Schools and Post Primary Schools.",
"descriptionIrish": "Not Disclosed",
"sourceUrl": "Not Disclosed",
"organization": "national-council-for-special-education",
"organizationTitle": "National Council for Special Education",
"organizationId": "46b1ba76-5e24-441e-91da-0d92e5c64616",
"organizationDescription": "The National Council for Special Education (NCSE) was set up to improve the delivery of education services to persons with special educational needs arising from disabilities with particular emphasis on children. The Council was first established as an independent statutory body by order of the Minister for Education and Science in December 2003.",
"theme": "Education and Sport",
"themeLabel": "Education and sport",
"tags": [
"2026-2027",
"SET",
"SNA",
"SNA Allocations",
"School Allocations"
],
"tagCount": 5,
"groups": [],
"licenceId": "CC-BY-4.0",
"licenceLabel": "Creative Commons Attribution 4.0",
"licenceUrl": "Not Disclosed",
"isOpenLicence": "No",
"isOpenDefinitionLicence": "Yes",
"rights": "Not Disclosed",
"isHighValueDataset": "No",
"hvdCategoryUri": "N/A",
"hvdCategory": "N/A",
"applicableLegislation": [],
"updateFrequency": "Daily",
"updateFrequencyLabel": "Every day",
"language": "en",
"version": "43",
"conformsTo": "Not Disclosed",
"provenance": "Not Disclosed",
"issued": "2026-06-09",
"updated": "2026-08-27",
"metadataCreated": "2026-06-09T16:40:10.095619",
"metadataModified": "2026-08-27T18:10:08.981851",
"temporalCoverage": "2026-2027 Academic Year",
"temporalStart": "Not Disclosed",
"temporalEnd": "Not Disclosed",
"spatialCoverage": "Not Disclosed",
"spatialUri": "Not Disclosed",
"spatialResolution": "Not Disclosed",
"boundingBox": "N/A",
"coordinateSystems": [],
"verticalExtentMin": "Not Disclosed",
"verticalExtentMax": "Not Disclosed",
"verticalDatum": "Not Disclosed",
"contactName": "NCSE Allocations",
"contactEmail": "allocations@ncse.ie",
"contactPhone": "Not Disclosed",
"author": "Not Disclosed",
"authorEmail": "Not Disclosed",
"maintainer": "Not Disclosed",
"maintainerEmail": "Not Disclosed",
"collectionName": "ncse-ckan",
"harvestSourceTitle": "National Council for Special Education",
"harvestSourceId": "43e35ce9-ae96-4e44-b26d-d43ff5913694",
"isHarvested": "Yes",
"resourceCount": 3,
"resourceFormats": [
"CSV"
],
"hasMachineReadable": "Yes",
"datastoreResourceCount": 3,
"totalSizeBytes": 291789,
"primaryDownloadUrl": "https://opendata.ncse.ie/dataset/735d61de-6400-4b81-832d-8ff6cc84fa4f/resource/f1f60760-d195-4be7-9a05-1a97159cbef6/download/sna-and-set-hour-allocations-primary-schools.csv",
"state": "active",
"rowType": "dataset",
"opennessScore": 3,
"opennessStars": "3 stars - structured data in a non-proprietary open format",
"opennessReason": "One of the resource formats is 3-star data - machine-readable data in an open format.",
"orgAcronym": "NCSE",
"orgDescription": "The National Council for Special Education (NCSE) was set up to improve the delivery of education services to persons with special educational needs arising from disabilities with particular emphasis on children. The Council was first established as an independent statutory body by order of the Minister for Education and Science in December 2003.",
"orgDatasetCount": 65,
"orgFollowerCount": 0,
"orgImageUrl": "https://data.gov.ie/uploads/group/2025-04-23-125021.575987ncse-logo.gif",
"orgCreated": "2025-03-03T10:14:28.330747",
"checkedDistributions": [
{
"resourceId": "f1f60760-d195-4be7-9a05-1a97159cbef6",
"url": "https://opendata.ncse.ie/dataset/735d61de-6400-4b81-832d-8ff6cc84fa4f/resource/f1f60760-d195-4be7-9a05-1a97159cbef6/download/sna-and-set-hour-allocations-primary-schools.csv",
"format": "CSV",
"alive": "Yes",
"status": 200,
"reason": "OK",
"contentType": "text/csv",
"contentLength": 219592
},
{
"resourceId": "46288931-2b8b-4dc8-907d-b18dacc471cd",
"url": "https://opendata.ncse.ie/dataset/735d61de-6400-4b81-832d-8ff6cc84fa4f/resource/46288931-2b8b-4dc8-907d-b18dacc471cd/download/sna-and-set-hour-allocations-post-primary-schools.csv",
"format": "CSV",
"alive": "Yes",
"status": 200,
"reason": "OK",
"contentType": "text/csv",
"contentLength": 61602
},
{
"resourceId": "c1acbfed-0adf-4322-b02b-52e390f09aef",
"url": "https://opendata.ncse.ie/dataset/735d61de-6400-4b81-832d-8ff6cc84fa4f/resource/c1acbfed-0adf-4322-b02b-52e390f09aef/download/allocations-special-schools.csv",
"format": "CSV",
"alive": "Yes",
"status": 200,
"reason": "OK",
"contentType": "text/csv",
"contentLength": 10595
}
],
"distributionsChecked": 3,
"distributionsAlive": 3,
"distributionsDead": 0,
"datastoreTables": [
{
"resourceId": "f1f60760-d195-4be7-9a05-1a97159cbef6",
"name": "SNA And SET Hours Allocations Primary Schools",
"rowCount": 3049,
"columnCount": 11,
"columns": [
"County",
"Dublin Area Codes",
"Roll Number",
"School Type",
"School Name",
"set_hours_26_27",
"set_posts__26_27",
"special_class_teaching_posts__26_27",
"mainstream_sna_allocation__26_27",
"special_class_snas__26_27",
"total_sna_allocation_26_27"
]
},
{
"resourceId": "46288931-2b8b-4dc8-907d-b18dacc471cd",
"name": "SNA And SET Hours Allocations Post Primary Schools",
"rowCount": 721,
"columnCount": 11,
"columns": [
"County",
"Dublin Area Codes",
"Roll Number",
"School Type",
"School Name",
"set_hours_26_27",
"set_posts__26_27",
"special_class_teaching_posts__26_27",
"mainstream_sna_allocation__26_27",
"special_class_snas__26_27",
"total_sna_allocation_26_27"
]
},
{
"resourceId": "c1acbfed-0adf-4322-b02b-52e390f09aef",
"name": "Allocations Special Schools",
"rowCount": 133,
"columnCount": 12,
"columns": [
"Roll Number",
"School Name",
"Admin Principal",
"Admin Deputy Principal",
"teaching_posts_26_27",
"exceptional_teaching_posts_26_27",
"Department of Education and Youth Concessionary Post",
"sna_posts_serc_26_27",
"additonal_sna_allocation__26_27",
"total_sna_posts_26_27",
"Cooperation Hours - Historic and New",
"Part-time Specialist Subject Hours"
]
}
],
"datastoreTablesProbed": 3,
"datastoreTotalRows": 3903,
"scrapedAt": "2026-08-27T19:58:06.625Z"
}

Fields the portal does not hold for a given dataset come back as Not Disclosed, fields that do not apply as N/A, and booleans as Yes or No. There are no literal nulls anywhere in the output.

Configure the run

Leave everything empty and the Actor sweeps the whole portal, newest change first. Add filters to narrow it, tick a row-type box to get extra rows, tick an enrichment box to get extra columns. Filters combine with AND; each list inside a filter combines with OR.

Everything from the Central Statistics Office published since the start of 2026, as CSV:

{
"organizations": ["central-statistics-office"],
"formats": ["CSV"],
"modifiedFrom": "2026-01-01",
"maxItems": 5000,
"sortBy": "metadata_modified desc"
}

The EU high-value compliance picture, with the publisher directory alongside it:

{
"onlyHighValue": true,
"includeOrganizations": true,
"includeHvdCategories": true,
"includeOpennessScore": true,
"maxItems": 800
}

Three named datasets with every enrichment, including what is actually inside their files:

{
"startUrls": [
{ "url": "https://data.gov.ie/dataset/planning-permission" },
{ "url": "https://data.gov.ie/dataset/valuation-office-api" }
],
"datasetSlugs": ["homelessness-report-june-2026"],
"includeResources": true,
"includeDatastoreProbe": true,
"includeLinkCheck": true,
"includeOpennessScore": true,
"datastorePreviewRows": 5
}

Pricing

Pay per event, so you are charged for the rows you actually receive and the extras you actually switch on. Nothing else.

EventPriceCharged when
Actor start$0.020Once per run
Dataset$0.007One dataset row
Portal page scanned$0.007One index page of up to 200 datasets
File$0.004One file row, only with the file rows box ticked
Publisher / Harvest source / Showcase / Data request$0.004One directory row of that type
Theme / EU category / Licence / Frequency / Collection / Format / Portal change$0.002One directory or change-feed row
Tag$0.001One tag row
Openness score attached$0.003Only when the portal returned a rating
DataStore table read$0.005Only when the table answered
File link checked$0.003Only when the link check ran
Publisher profile attached$0.003Only when the profile came back

Bronze, Silver and Gold plans pay 3.8%, 7.4% and 11% less per event.

Datasets, metadata onlyEvents chargedCost
1001 start, 100 datasets, 1 page$0.73
1,0001 start, 1,000 datasets, 5 pages$7.06
10,0001 start, 10,000 datasets, 50 pages$70.37

Free users

Apify free-plan accounts get 10 rows per run as a preview, enough to see every field and check the shape before you commit. Paid plans return up to 1,000,000 rows per run. Upgrade here.

Run it

  1. Create a free Apify account. New accounts get $5 in platform credit, which covers about 700 dataset rows here.
  2. Open the Actor, leave the input as it is for a 10-row preview, or paste one of the examples above.
  3. Press Start. A 10-row preview finishes in about 4 seconds, a metadata-only sweep wrote 1,679 dataset rows in 5 seconds, and 1,531 datasets with every enrichment block on took 575 seconds.
  4. Download the dataset as CSV, JSON, Excel or XML, or read it from the Apify API.

Use with AI agents (MCP)

claude mcp add apify npx -- -y @apify/actors-mcp-server --actors parseforge/data-gov-ie-ireland-open-data-scraper

Then ask in plain language:

  • "List every dataset the Environmental Protection Agency published on data.gov.ie since 2025, with its licence and file formats."
  • "Which Irish public bodies publish EU high-value datasets, and how many each?"
  • "Check the planning permission dataset on data.gov.ie: are its files still downloadable, and what columns does the CSV have?"

Troubleshooting

I get no results at all. A filter value that does not exist upstream returns zero rows, not everything. Check the spelling of a publisher name or a tag against the portal, or tick the matching directory box to pull the real vocabulary first.

I get fewer rows than I asked for. maxItems is one budget shared by every row type in the run. Datasets and file rows are written first, then the directories smallest first, and the tag directory last, so a low maxItems with many boxes ticked spends itself before it reaches the tags. Raise maxItems, or run the directories in their own run.

My format filter finds nothing. The portal's index is case sensitive: res_format:csv matches zero datasets and res_format:CSV matches 20,237. The format list in the input already expands each choice to every spelling the portal uses, so pick from the list rather than typing a format in the custom Solr filter.

A field I expected is "Not Disclosed". That means the publisher did not supply it. Nine tenths of the portal is harvested from 31 external feeds and each one fills a different subset: frequency is present on 36% of datasets, rights on 7%, a spatial geometry on 6%. N/A means the field does not apply to that row type at all.

The run is slower than I expected. Metadata-only paging is fast: 1,679 dataset rows landed in 5 seconds on a measured run. Each enrichment block adds a network round trip per dataset or per file, so a run with the openness score, the link check, the DataStore probe and the publisher profile all switched on lands nearer 14 rows per second. Turn off the blocks you do not need.

FAQ

QuestionAnswer
Do I need a data.gov.ie account or API key?No. The portal's CKAN API is fully anonymous and this Actor uses no proxy.
How many datasets are on data.gov.ie?22,666 as measured on 2026-08-27, across 174 publishing bodies and 14 themes.
Why does the tag directory stop at 2,000 rows?The portal returns at most 2,000 values per facet, so the tag directory is the 2,000 most used tags out of a vocabulary of 22,742. The run logs a warning when it hits that ceiling.
Can I get the actual data inside a file, not just the metadata?For the 19.4% of files loaded into the portal DataStore, yes: tick the DataStore probe for column names, SQL types, row counts and sample rows. For the rest you get the direct download URL.
What is the openness score?The portal's own 0 to 5 star rating of how open a dataset is, following the five-star open data model. It is on every dataset but only reachable through the record view, so it is an opt-in block here.
Why does a CC-BY-4.0 dataset say isOpenLicence: No?Because the portal says so. 732 CC-BY-4.0 datasets are flagged not open upstream. The row carries both the portal's flag and an independent Open Definition check in isOpenDefinitionLicence.
Can I filter by map area or bounding box?Not by box. The portal has the spatial search extension installed but it answers HTTP 409 on any bounding box, so this Actor offers onlyGeospatial plus the spatial coverage name and computed bounding box on each row instead.
How do I monitor the portal for changes?Tick the change feed box and run on a schedule. Each row is one dataset change with its timestamp, so you can diff against the previous run and refetch only what moved.
Is the data in Irish as well as English?Rarely. 22,665 datasets are tagged English and one Irish. 373 of 400 sampled datasets carry a translation field, but only 20 hold an Irish title that actually differs from the English, and only those are shipped in titleIrish.
What happens if a dataset I ask for does not exist?The run writes a single row with rowType: "error" naming the URL, and carries on with the rest. Error rows are never charged.
Can I sort by anything I like?Only by the thirteen orders in the dropdown. The portal silently falls back to relevance for anything else, so an unsupported sort is rejected with a warning rather than quietly ignored.

Browse the full ParseForge collection for more scrapers.

Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

Disclaimer: this is an unofficial scraper and is not affiliated with, endorsed by or connected to data.gov.ie, the Department of Public Expenditure, NDP Delivery and Reform, or any Irish public body. It reads only data the portal publishes anonymously to anyone. Dataset metadata is published under the licences named in each row, most commonly Creative Commons Attribution 4.0, and you are responsible for honouring the attribution those licences require. Contact fields in the output are the published contact points of public bodies rather than personal data, but if a record does carry personal information you remain the controller of what you collect and must handle it in line with GDPR, CCPA, PIPL and any other law that applies to you.