dane.gov.pl Scraper - Poland Open Data Portal API avatar

dane.gov.pl Scraper - Poland Open Data Portal API

Pricing

from $6.23 / 1,000 datasets

Go to Apify Store
dane.gov.pl Scraper - Poland Open Data Portal API

dane.gov.pl Scraper - Poland Open Data Portal API

Scrape dane.gov.pl, Poland's national open data portal: 26,922 datasets, 1.6M resources, 7,471 institutions, showcases, change history and table rows over the public JSON:API.

Pricing

from $6.23 / 1,000 datasets

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

ParseForge

dane.gov.pl Scraper - Poland Open Data Portal API

Pull all 26,989 datasets, 1.6 million resources and 7,476 publishing institutions out of dane.gov.pl, Poland's national open data portal, in one run. Every dataset row carries 49 fields: Polish title and description, publisher with its REGON and address, EU DCAT themes, licence and every reuse condition, update frequency, openness score, file formats, view and download counts and four timestamps. No login, no API key, no rate limit. Exports to CSV, JSON, Excel and XML.

The portal has a public JSON:API, and that sounds like it settles the matter until you try to use it. The listing endpoint ignores a filter name it does not know and hands back the entire catalogue with HTTP 200, so a typo looks like a successful unfiltered query. sort behaves the same way. The tag filter is in the official spec and returns HTTP 400. Licence and update frequency filter on the search index and nowhere else. The resource index stops paging at exactly 40,000 rows with all shards failed, the audit log at 10,000. Single requests hang for a full minute at random. This scraper knows all of that, measures every filter against the unfiltered total before it sends it, paginates around the walls, and gives you a flat table instead of nested JSON:API envelopes.

Who uses itWhat they scrape dane.gov.pl for
Data engineersBuilding an open data catalog mirror with formats, licences and refresh cadence per dataset
Property analystsThe 6,635 developers legally obliged to publish flat price lists, and their daily updates
Compliance teamsHigh value datasets under the Polish open data act and the EU HVD list
Civic tech and journalistsWhich public bodies publish what, how often, and which of their links are already dead
AI and RAG buildersPolish public sector text and real table rows as grounded training and retrieval material

What it does

  • 📚 Ten collections in one Actor - datasets, resources, institutions, showcases, change history, the broken links report, news, knowledge base, cross-model search, and the actual rows inside a table.
  • 🇵🇱 49 fields on every dataset - title and description in plain text and HTML, publisher with type, city and website, DCAT themes with EU codes, keywords, regions, formats, delivery types, licence, reuse conditions, openness score, update frequency, views, downloads, harvest source and four timestamps.
  • 📊 Real table rows, with real headers - the data API returns col1, col2, col3; this Actor restores the publisher's own column names, so you get Nazwisko aktualne and Liczba, not numbered placeholders.
  • 🔗 Live link checking - one HEAD request per resource records the HTTP status, content type and byte size of the file itself, which finds dead downloads the portal has not flagged yet.
  • 🏛️ Full publisher profiles - REGON, postal address, email, telephone, fax, ePUAP inbox, electronic delivery address, website and the external catalogues the body is harvested from.
  • 🧮 Catalogue breakdowns - one row per format, theme, category, delivery type, openness score and publisher, with counts and share of whatever slice you filtered to.
  • 🇬🇧 English or Polish - English translates the controlled vocabularies and links to the English portal. Titles and descriptions are written in Polish by the publishers, so they stay Polish either way.
  • ⏱️ Paging that respects the walls - the resource index dies past 40,000 rows, the audit log and any single table past 10,000; the run stops cleanly at each ceiling and logs exactly how many rows are out of reach instead of failing. A page that times out is retried, never dropped in silence.

What you can do with dane.gov.pl data

Mirror Poland's open data catalog. Pull all 26,989 datasets with their formats, licences, openness scores and update frequencies, then re-run weekly with a Modified from date to catch only what changed. The 49-field row is flat, so it lands in a warehouse table without any JSON flattening step.

Track the property market at source. Polish developers must publish their flat price lists as open data under the 2021 developer act, which is why 6,635 of the 7,476 registered institutions are developers and why 21,858 of the datasets sit under Economy and finance. Filter to Property developer, turn on the file check and the table preview, and you get the price lists with their live download URLs.

Audit a public body's transparency. Point it at one institution ID, turn on the change history block, and you get every dataset that body publishes, when each was last refreshed against its promised cadence, and the full audit trail of who edited what.

Feed a RAG index with Polish public sector text. Datasets, news and knowledge base articles all carry full descriptions, and the table rows collection gives you the underlying data itself. It is public domain or CC BY in almost every case, so the licence question is already answered in the row.

Why choose this scraper

What you get
Every filter measuredEach filter was sent against the unfiltered total before it shipped. The ones this API accepts and then silently ignores are applied after the fetch instead, on the full row, before anything is billed
Nothing always emptyThe five sparse licence-condition columns are folded into one readable list, and no field ever comes back as a literal null
The walls documented40,000 rows on resources, 10,000 on the audit log, 10,000 per table on the data API, 26,989 datasets fully reachable, each one bisected and logged at run time
Ten record typesNot just the catalogue: the files, the publishers, the reuse showcases, the audit log, the dead-link report, the editorial content and the table contents
Pay for what returnedAn enrichment block that came back empty is not charged, and the run asks the charging manager before it writes a batch rather than after
No proxy, no keyThe API is open. There is no proxy cost in your bill and no credential to manage

How it compares

There is one other dane.gov.pl Actor in the Apify Store, benthepythondev/poland-dane-gov-scraper, at $2.50 per 1,000 rows against our $7.00. It is cheaper and it has two billing events against our nineteen, which is the honest summary of the difference: it charges for a start and a row, and that is what it does. Ours costs 2.8x more per dataset row and in exchange covers ten record types instead of one, carries 49 fields on the primary row, restores real column headers on table data, checks whether files still download, and knows where each index stops paging. If you want one flat pass over the dataset list and nothing else, the cheaper one is the right call.

This Actorbenthepythondev/poland-dane-gov-scraperThe portal's own API
Price per 1,000 datasets$7.00$2.50Free
Billing events192n/a
Record types10120 endpoints, JSON:API envelopes
Fields on a dataset row49not published37 raw attributes, nested
Table rows with real headersYesNocol1, col2, col3
Live file checkingYesNoNo
Paging walls handledYesNot publishedFails with all shards failed
Users in the last 30 daysnew0n/a

What a dane.gov.pl dataset looks like

One unedited row from the verified platform run fW4CEt1RpHtSsx2t1:

{
"datasetId": "1",
"slug": "dane-liczbowe-dot-kontroli-prowadzonych-przez-wiih-w-2014-r",
"title": "Dane liczbowe dot. kontroli prowadzonych przez WIIH w 2014 r.",
"description": "Dane liczbowe dot. kontroli prowadzonych przez wojewódzkich inspektorów Inspekcji Handlowej (IH) w 2014 r.",
"descriptionHtml": "<p>Dane liczbowe dot. kontroli prowadzonych przez wojewódzkich inspektorów Inspekcji Handlowej (IH) w 2014 r.</p>",
"url": "https://dane.gov.pl/pl/dataset/1,dane-liczbowe-dot-kontroli-prowadzonych-przez-wiih-w-2014-r",
"apiUrl": "https://api.dane.gov.pl/1.4/datasets/1,dane-liczbowe-dot-kontroli-prowadzonych-przez-wiih-w-2014-r",
"externalUrl": "Not Disclosed",
"institutionId": "26",
"institutionName": "Urząd Ochrony Konkurencji i Konsumentów",
"institutionType": "Administracja rządowa",
"institutionCity": "Warszawa",
"institutionWebsite": "https://uokik.gov.pl/index.php",
"dcatThemes": [{ "id": "148", "code": "GOVE", "title": "Rząd i sektor publiczny" }],
"dcatThemeCodes": [
"GOVE"
],
"legacyCategory": "Administracja Publiczna",
"keywords": [
"uokik",
"inspekcja handlowa"
],
"regions": [
"Polska"
],
"formats": [
"xlsx"
],
"resourceTypes": [
"file"
],
"visualizationTypes": [],
"resourceCount": 1,
"showcaseCount": 0,
"licenseName": "CC0 1.0",
"licenseCode": "N/A",
"licenseDescription": "Other (Public Domain)",
"licenseConditions": [],
"updateFrequency": "Rocznie",
"opennessScore": 2,
"opennessScores": [
2
],
"hasHighValueData": "No",
"hasEuHighValueData": "Not Disclosed",
"hasDynamicData": "No",
"hasResearchData": "No",
"isPromoted": "No",
"viewsCount": 673,
"downloadsCount": 138,
"harvestedFromTitle": "Not Disclosed",
"harvestedFromUrl": "Not Disclosed",
"harvestedFromType": "Not Disclosed",
"harvestedLastImport": "Not Disclosed",
"supplements": [],
"imageUrl": "Not Disclosed",
"imageAlt": "Not Disclosed",
"createdAt": "2016-04-19T07:39:49Z",
"modifiedAt": "2024-03-19T19:14:09Z",
"verifiedAt": "2016-04-19T09:41:11Z",
"resourceModifiedAt": "2022-12-05T11:40:46Z",
"scrapedAt": "2026-08-27T20:21:21.768Z"
}

Configure the run

Pick one or more collections, set how many rows you want, and add filters. Max Items caps the whole run, and with several collections selected each gets an even share of it, with anything one collection does not use rolling forward to the next. Everything else has a working default. The catalogue filters run inside the portal index; the post-fetch refinements are the fields the index accepts and then ignores, so they cost more upstream reads and are labelled as such in the schema.

Fresh CSV datasets from central government, newest first:

{
"collections": ["datasets"],
"formats": ["csv"],
"minOpennessScore": "3",
"createdFrom": "2026-01-01",
"sortBy": "-created",
"maxItems": 5000
}

Every file a single publisher offers, with a live download check and the table schema:

{
"collections": ["resources"],
"institutionIds": ["26"],
"includeTabularPreview": true,
"includeFileProbe": true,
"maxRowsPerResource": 5,
"maxItems": 2000
}

The most common Polish surnames, straight out of the PESEL register table:

{
"collections": ["tabular-rows"],
"resourceIds": ["28052"],
"maxRowsPerResource": 1000,
"maxItems": 1000
}

That table holds 298,692 rows and the data API pages 10,000 of them, so ask for more than that and the run stops at the ceiling and says so.

Pricing

Pay per event. You are charged for the rows you receive, plus a small fee per catalogue page paged through, plus one Actor start.

EventPriceWhen
Actor start$0.020Once per run
Catalogue page scanned$0.004Per page of up to 50 records
Dataset$0.007Per dataset row
Resource$0.003Per file or API row
Institution$0.008Per publisher row
Showcase$0.006Per app or service row
Change history entry$0.003Per audit log row
Broken link$0.004Per dead-link row
News item$0.004Per news row
Knowledge base article$0.004Per article row
Search result$0.003Per cross-model hit
Table row$0.002Per row of real data
Catalogue breakdown$0.002Per facet bucket
Files of a dataset$0.004Opt in, only when files came back
Apps reusing a dataset$0.004Opt in, only when a showcase exists
Change history of a dataset$0.004Opt in, only when the log had entries
Full publisher profile$0.005Opt in, only when the profile returned
Table preview$0.004Opt in, only for tabulated resources
Live file check$0.003Opt in, only when the check completed

What a plain dataset run actually costs:

DatasetsStartPagesRowsTotal
100$0.02$0.008$0.70$0.73
1,000$0.02$0.080$7.00$7.10
10,000$0.02$0.800$70.00$70.82

Bronze, Silver and Gold plans get 3.8%, 7.4% and 11% off every event.

Free users

Free Apify accounts get 10 rows per run as a preview, enough to check the field set and the filters before committing. Paid plans lift the cap to 1,000,000 rows. Upgrade here.

Run it

  1. Create a free Apify account. It comes with $5 of platform credit, no card needed.
  2. Open the Actor and pick your collections. Leave everything else alone for a first run.
  3. Set Max Items and press Start.
  4. Export from the Storage tab as CSV, JSON, Excel or XML, or pull it over the API.

Use with AI agents (MCP)

claude mcp add apify --transport sse https://mcp.apify.com/sse --header "Authorization: Bearer YOUR_APIFY_TOKEN"

Then ask in plain language:

  • "Get me every CSV dataset published on dane.gov.pl since January, with its publisher and licence."
  • "Which Polish public institutions have broken download links on the open data portal right now?"
  • "Pull the 500 most common Polish surnames out of the PESEL register dataset."

Troubleshooting

No results at all. The most likely cause is a format filter in the wrong case. This index stores formats lower case and CSV returns zero while csv returns 10,698. Narrow date ranges do the same thing quietly.

Fewer rows than I asked for. Three ceilings are real and all three are logged: the resource index stops at 40,000 rows, the audit log at 10,000, and a single table on the data API at 10,000, after which the portal answers all shards failed. Flip Sort by to reach the other end of an index, add a filter to cut the result set below the wall, or use Table row search to reach deeper rows of one table.

A field I expected is empty. Almost every field on this portal is optional. harvestedFrom* is filled on 39% of datasets, externalUrl on 31%, the legacy category on 17%. Empty means the publisher left it blank, and the row says Not Disclosed rather than hiding it.

My licence or update frequency filter did nothing. Those two only filter on the cross-model search index. In the dataset collection they are applied after the fetch, under Post-fetch refinements, which is slower but correct. Use the Search collection with Search: licence for the server-side version.

The run is slower than I expected. Single requests to this API hang for up to a minute at random, roughly one in five. The Actor times each attempt out at 25 seconds and retries against the clock, and runs six pages in parallel, which is why a 5,000 dataset run measured 63 rows per second even though one request averages three seconds. Turning on the opt-in blocks costs one extra call per row and drops that to around 8 rows per second.

FAQ

QuestionAnswer
Do I need an API key for dane.gov.pl?No. The portal's JSON:API is fully anonymous and this Actor uses no proxy and no credentials.
How many datasets are there?26,989 datasets, 1,613,952 resources, 7,476 institutions, 119 showcases and 269,811 audit log entries when this was measured. The catalogue grows daily, so treat these as a floor.
Can I download the actual CSV files?The row gives you the direct downloadUrl for each resource, and the table rows collection returns the parsed contents of tabulated resources without you downloading anything.
Why do so many publishers look like property companies?Polish developers are legally required to publish flat price lists as open data, so 6,635 of the 7,476 institutions are developers. Filter Institution type to Central government for the ministries.
Is the data in English?The controlled vocabularies are, when you set Language to English. Titles and descriptions are written in Polish by the publishers and are never translated by the portal.
Can I get only what changed since last week?Yes. Use Modified from in the dataset collection, or the Change history collection with a Changed from date for the audit trail itself.
What licence is the data under?24,098 datasets are CC0 1.0 and 2,825 are CC BY 4.0, with 66 across the NonCommercial and ShareAlike variants. Each row carries the licence name and its reuse conditions.
How fast is it?63 rows per second, measured on a 5,000 dataset run on the Apify platform that finished in 79 seconds. At that rate the whole catalogue takes about seven minutes. The opt-in blocks are one HTTP call each and cut it to roughly 8 rows per second.
What is the openness score?The portal's 1 to 5 star rating on Tim Berners-Lee's scale. 3 means a non-proprietary format like CSV, 4 adds URIs, 5 adds links out to other data. 13,990 datasets are 3 stars or better.
Does it work on the English portal?Yes. Setting Language to English switches the vocabularies and points every url at dane.gov.pl/en/....

Browse the full ParseForge collection for more scrapers.

Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

Disclaimer: this is an unofficial scraper, not affiliated with or endorsed by the Ministry of Digital Affairs, dane.gov.pl or the Polish government. It reads only data the portal publishes anonymously to anyone. Institution records contain business contact details of public bodies, published by those bodies as open data; handle any personal data you extract in line with GDPR, CCPA and PIPL, and keep the licence terms carried in every row.