Poland Pharmacy Register Scraper - Licensed Pharmacies + NIP avatar

Poland Pharmacy Register Scraper - Licensed Pharmacies + NIP

Pricing

from $2.50 / 1,000 per pharmacy returneds

Go to Apify Store
Poland Pharmacy Register Scraper - Licensed Pharmacies + NIP

Poland Pharmacy Register Scraper - Licensed Pharmacies + NIP

From $2.50 per 1,000 rows. Scrape Poland's official pharmacy register (Rejestr Aptek): all 23,993 licensed pharmacies with owner company, NIP, REGON, KRS, permit number and date, phone and e-mail. Canonical voivodeship, powiat, gmina, TERYT and opening hours joined from the state's own bulk export.

Pricing

from $2.50 / 1,000 per pharmacy returneds

Rating

0.0

(0)

Developer

Scrapers Delight

Scrapers Delight

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Share

💊 Poland Pharmacy Register Scraper — every licensed pharmacy and the company that owns it

Scrapes Rejestr Aptek, Poland's official register of pharmacy licences — the Krajowy Rejestr Zezwoleń na Prowadzenie Aptek Ogólnodostępnych, Punktów Aptecznych oraz Rejestr Udzielonych Zgód na Prowadzenie Aptek Szpitalnych i Zakładowych, kept by the 16 provincial pharmaceutical inspectorates and published by CeZ at rejestry.ezdrowie.gov.pl.

23,993 pharmacies. 13,733 of them open. Each row carries the pharmacy, its address with a canonical voivodeship, its phone and e-mail, the owner entity with NIP, REGON and KRS, and the permit — number, type, issuing inspectorate and issue date.

Every number on this page was measured against the live register on 2026-09-07 / 2026-09-08, over the whole 23,993-row corpus — not over a sample, and not estimated. Where a figure differs between the whole register and the live pharmacies only, both are shown, each labelled with the set it is over.


📊 What you get, per row

106 columns. Fill rates below are over all 23,993 rows and over the 13,733 active ones.

Identity and status

FieldCorpusActive
register_id — the state's own id, joins to the official export100%100%
name — trading name83.8%88.9%
pharmacy_type + _code + _en — 10 kinds100%100%
status + _code + _en + is_active — 5 statuses100%100%
registration_number39.8%40.2%

16.2% of the register genuinely has no trading name — mostly hospital pharmacy departments, which are recorded under the hospital rather than under a shop sign.

Contact — this is a dialable list, not a list you then have to enrich

FieldCorpusActive
email83.5% (20,045)98.4% (13,520)
phone89.5% (21,477)97.8% (13,427)
phone or e-mail91.1% (21,866)99.4% (13,652)
fax9.9%10.5%
website1.6%2.1%
sells_online (mail-order flag)1.5% (354)1.8% (247)

No enrichment hop, no e-mail-guessing vendor, no per-contact surcharge. These are the contacts the pharmacy filed with the state. website is effectively absent — 1.6% — and we are not going to pretend otherwise.

Address, with the region actually usable

FieldCorpusActive
city, powiat, voivodeship100%100%
postcode99.9%100%
gmina98.5%100%
street, house_number, flat_number, address_line91.4% street93.4%
terc, simc, ulic — TERYT statistical codes50.2%52.1%
latitude / longitude — WGS84 degrees, XML export only21.9% (5,258)34.5% (4,737)
puwg92_x / puwg92_y — the raw grid pair the conversion came from21.9%34.5%
voivodeship_raw + voivodeship_source — auditability100%100%

Owner entity — where the independents separate from the chains

FieldCorpusActive
owner_nip — Polish tax number93.0%99.8% (13,702)
owner_regon93.5%99.9% (13,724)
owner_krs — company-court number54.7%62.7% (8,604)
owner_name90.9%98.2%
owner_legal_form + _code + _id + _en — 20 forms96.7%100%
owner_is_independent / owner_is_company / owner_has_krs100%100%
owner_first_name, owner_last_name — sole traders35.8%31.4%
Owner address: street, city, postcode, owner_voivodeship, owner_powiat, owner_gmina, owner_terc, owner_simc96.6% city100%

7,225 distinct owner NIPs run the 13,733 active pharmacies. 2,120 of those owners run more than one; the largest single owner runs 101. And 4,911 active pharmacies (35.8%) belong to an independent — a sole trader, a civil-law partnership or another non-registered person. One tick of ownerType: "independent" is the whole independents list.

Permit

FieldCorpusActive
permit_number, permit_type, permit_issuer (71 authorities)98.7%98.4%
permit_issue_date + permit_age_days98.6%98.4%
permit_effective_date91.8%89.8%
permit_changes_count + last change date / document / description44.8%52.3%
permit_history[] — every amendment (opt-in)44.8%52.3%
permit_revoke_date, permit_withdrawal_date, permit_suspension_date, permit_expiry_date0.2%–35.7%~0%

permit_expiry_date is filled on 40 of 23,993 rows (0.2%). It is in the schema for completeness and it is not a feature. Polish pharmacy licences are open-ended.

Operations (from the official bulk export)

launch_date (99.0% / 98.9%) · opening_hours per weekday (66.6% corpus, 90.8% of active) · activity_scope (0.1% — 20 rows) · manager_first_name, manager_last_name, manager_licence_number_pwz, manager_since, and the deputy manager (61.7% corpus, 99.9% of activeopt-in, personal data, see below) · suspension_date and restart_date (XML export).

Provenance on every row

source_url (a one-row lookup you can paste into a browser) · scraped_at · export_source and export_snapshot_date (the register's own stan na dzień) · export_joined. Plus a RUN_SUMMARY record in the key-value store: delivered vs declared counts, per-stream reconciliation, what each filter dropped, and — if the run was cut short — which limit did it.

Two flags in that record are the ones worth automating against:

FieldTrue only when
register_completethe dataset holds every row that matched your query — nothing truncated, nothing empty, nothing held back. register_complete_reason names the shortfall whenever it is false.
read_completethe crawl itself paged to the end of every stream and reconciled its unique register ids against the register's own declared count.

They are deliberately two different questions. A run can read a result set perfectly and still deliver you nothing, so register_complete describes what you received, not what was fetched. It is false on a zero-row run, whatever the reason — an empty dataset is never a complete snapshot.

It is also not false for the wrong reason. A run that delivers exactly as many rows as your Max pharmacies setting allows, when that happens to be every row your query matched, is complete and says so: the cap counts against you only when it actually refused a row the crawl had in hand. Measured — city: "Opole", maxItems: 113, against a query the register answers with 113: 113 fetched, 113 delivered, 0 dropped, register_complete: true.

You do not have to open the key-value store to see any of this. The run's status message carries the same verdict: rows delivered, rows billed, the limit that stopped the run if one did, and whether what you hold is the complete matching set.


🎯 Who buys this

  • Pharmaceutical wholesalers and distributors (Neuca, Farmacol, Salus and their reps) — a territory list with a canonical voivodeship, a powiat, a phone and an e-mail on 99.4% of live pharmacies, and the owner NIP to match against the account book.
  • Franchise and buying-group recruitersownerType: "independent" returns the 4,911 active independents, exactly the accounts that are not already inside a chain.
  • Pharmacy software, POS, e-prescription and shelf-labelling vendors — the whole addressable market of Polish pharmacies with a dialable contact, segmented by region and by chain size.
  • Market analysts and PE / M&A — 7,225 owning entities, their legal form, their KRS, and a permit-date series that shows openings and closures.
  • Anyone joining to Polish company data — NIP on 99.8% and REGON on 99.9% of active rows means a clean key into KRS, CEIDG, REGON/BIR and any credit-scoring source, with no matching step.

🔎 What you can filter on — and where the filter runs

The register honours 15 query parameters and silently ignores every other one, answering HTTP 200 with the entire 23,993-row corpus. This Actor therefore keeps an allowlist, refuses to send anything else, and asserts that a filtered stream's declared total is below the unfiltered total — so a run can never quietly bill you for the whole register in place of your filter.

Server-side (cheap — the register does the work): city · powiat · gmina · street · pharmacy name · owner name · owner first/last name · permit number · permit issue date, from and to · owner legal form · voivodeship · register id · registration number.

Client-side (the register genuinely cannot do these — the rows are fetched and then dropped before delivery, so you are never charged for them): status / active-only · pharmacy type · permit type · permit issuer · owner NIP · "must have e-mail / phone" · sells-online · postcode prefix · record-change window.

⚠️ Text search is case-insensitive but DIACRITIC-SENSITIVE. Kraków → 531 pharmacies, krak → 531, krakow → 0. Type the Polish spelling. The Actor does not silently fold your term, because that would change what you asked for.

⚠️ A voivodeship filter forces the per-voivodeship crawl, and says so. pharmacyProvince is a per-request server filter, so only a per-region stream can carry it — and address.province is null on 100% of the search endpoint's rows, so a single national stream has no way to tell which region a row is in. If you pick regions and set the crawl strategy to single-stream, the Actor overrides you to by-voivodeship and logs why, rather than crawling the whole country and billing you for it. Belt and braces: the client-side region check fails closed — a row whose region this run could not establish is dropped, never delivered unverified. (If you combine explicit register ids with a voivodeship filter and turn the export join off, that means every row is dropped; the log warns you up front.)


🚀 Example inputs

The whole live register, region-ready

{ "maxItems": 0, "activeOnly": true, "includeOfficialExportFields": true }

13,733 rows. Measured end to end at 94.6 s for the full 23,993-row read (16 parallel voivodeship streams, 33 page requests, 18.8 s of that the bulk export).

Independent pharmacies in Masovia with a phone or an e-mail

{ "voivodeships": ["mazowieckie"], "ownerType": "independent",
"requireEmailOrPhone": true, "activeOnly": true, "maxItems": 0 }

Newly licensed pharmacies — the real trigger event

{ "sortBy": "permission.issueDate", "sortDirection": "DESC",
"permitIssuedInLastDays": 90, "activeOnly": false, "maxItems": 0 }

Be realistic about the size of this signal: 12 permits in the last 30 days, 46 in 90, 200 in 365. Of the 46 in the last 90 days, 40 are already open and 6 are still oczekująca — permit granted, pharmacy not yet trading, which is the highest-intent row in the file.

One chain's whole estate, by tax number

{ "ownerNip": "9512323159", "activeOnly": false, "maxItems": 0 }

Monitor: what the register changed this week

{ "changedInLastDays": 7, "dedupeAcrossRuns": true, "maxItems": 0 }

38 records changed in the last day, 381 in 7 days, 652 in 30, 1,862 in 90.

Look up specific pharmacies by register id (one exact request each, no crawl)

{ "registerIds": ["1232429", "1000015"], "activeOnly": false }

💵 Pricing

EventPriceWhen it fires
Per pharmacy returned$0.0025 ($2.50 / 1,000)Once per row actually written to your dataset
Official bulk-export join$0.02, once per runOnly after the export downloaded and at least one delivered row gained a field from it

No actor-start charge. No per-result auto-event. No monthly rental.

RunRowsCost
The whole register23,993$59.98 + $0.02
Active pharmacies only (the default scope)13,733$34.35
One voivodeship, active only (e.g. Masovia)~1,750~$4.40
Active independents nationwide4,911~$12.30
Newly licensed, last 90 days46~$0.14

You are never charged for a row you did not receive. Duplicates are removed by register id before billing. Rows dropped by a client-side filter are never delivered and never charged. Rows skipped because an earlier run already delivered them are never charged. If a charging limit (maxTotalChargeUsd, or a free-tier balance) truncates a run, the batch is trimmed before the push and the log names the limit as the cause instead of blaming the register.


⚙️ How it works, and what it does not hide

Plain HTTP against the register's own public JSON endpoint. No key, no cookie, no session, no CAPTCHA, no geo-block, no browser. There is no robots.txt to quote: both hosts answer /robots.txt with HTTP 200 and the Angular application shell, not a robots file.

GET /api/ra/pharmacies/search?page=N&size=N&sortField=…&sortDirection=…[&filters]
-> {"23993":[ …rows ]}

Four things this register does that will silently ruin a naive scrape

1. sortField and sortDirection are mandatory — and undocumented. ?page=0&size=5 returns HTTP 500. So does ?page=0&size=5&sortField=dateOfChanged. Only the full pair returns 200. A scraper that omits either fails 100% of the time.

2. The row count is the JSON key, not a field. The body is {"23993":[…]}. There is no total or content wrapper. That key is also a free oracle: if a filter you sent comes back with the unfiltered total, the server ignored it.

3. Offset paging over the default sort loses rows. Measured on małopolskie (2,171 declared), crawled contiguously:

Sort keyFetchedUniqueDuplicates
originId (register id), page size 1,0002,1712,1710
originId, page size 2002,1712,1710
dateOfChanged DESC, page size 1,0002,1712,1656
dateOfChanged DESC, page size 2002,1712,1638

The rows the change-sorted crawl duplicated are rows it missed: the register-id crawl returned 6 ids the change-sorted crawl never produced at all. dateOfChanged is null on 20.1% of the corpus and is rewritten by bulk admin edits, so ties reshuffle between requests. Register id is unique, immutable and monotonic — so this Actor sorts by it by default, and warns in the log when you choose a mutable key.

Proven at full scale, three independent crawls:

CrawlDeclaredFetchedUniqueWall
register id ASC, page 1,00023,99323,99323,993140.6 s
register id DESC, page 1,00023,99323,99323,993144.8 s
register id DESC, page 50023,99323,99323,993186.1 s

All three id sets are identical to each other and identical to the official bulk export's 23,993 ids — zero missing, zero extra.

4. Whole columns exist in the response and are never filled. Over all 23,993 search rows, address.province, address.commune, address.district, address.latitude, address.terc, address.simc, workingHours, managersInfo and activityScopes are null on 100% of them, and the eight boolean flags (wholeDayPharmacy, temporaryClosed, internPharmacy, asepticCondition, …) are false on all 23,993. Those flags are deliberately not emitted — shipping a column that is false everywhere is worse than shipping nothing. The rest come from the official bulk export.

The voivodeship problem — the actual work in this Actor

The register's JSON returns address.province null on 100% of rows, and the state's own bulk export spells the 16 voivodeships 50 different ways: MAZOWIECKIE / Mazowieckie / mazowieckie, Kujawsko pomorskie vs kujawsko-pomorskie, Warmińsko mazurskie vs warmińsko-mazurskie, and so on. Scrape either route naively and you ship rows that cannot be grouped or filtered by region — which is the first thing a field-sales manager does.

This Actor stamps the canonical name from the register's own server-side pharmacyProvince partition, and normalises the export's spelling to the same 16 values. Both were verified lossless and verified to agree, value for value:

mazowieckie3,060podkarpackie1,369
śląskie2,843pomorskie1,306
łódzkie2,268kujawsko-pomorskie1,176
wielkopolskie2,173zachodniopomorskie1,004
małopolskie2,171świętokrzyskie789
dolnośląskie1,763warmińsko-mazurskie769
lubelskie1,390podlaskie719
opolskie615
lubuskie578

Sum: 23,993 — exactly the corpus, nothing double-counted, nothing lost. The original spelling is kept in voivodeship_raw and the provenance in voivodeship_source, so you can audit the mapping.

The dead-row trap

42.8% of the register — 10,260 of 23,993 rows — is closed, suspended or temporarily shut.

StatusRows
aktywna (active)13,733
nieaktywna (closed)9,950
nieaktywna – zawieszenie działalności (suspended)166
czasowo nieczynna (temporarily closed)124
oczekująca (permit granted, not yet open)20

A scrape of page 0 sees 84% active and reports 98.9% e-mail fill. The truth over the whole corpus is 57.2% active and 83.5% e-mail. That is exactly why activeOnly is on by default and why both sets of fill rates are published on this page. Load the register without this filter and you put 10,260 dead pharmacies into a CRM.

Do not mistake the change feed for a new-pharmacy feed

dateOfChanged is administrative churn. Of the 1,000 most recently changed records, 94.9% carry a permit more than a year old, the median permit age is 15.5 years, and only 1.2% were licensed in the last 30 days. 201 records were bulk-edited on 2026-09-01 alone. The honest new-pharmacy feed is sortBy: "permission.issueDate" with permitIssuedInLastDays — a genuine server-side date range.


📡 Reliability and speed, measured on the Apify platform

Every row below is this Actor running in an Apify container, not on a laptop.

RunRowsTime
Prefill demo — active pharmacies + the export join507.4 – 26.9 s (5 runs; see the note below)
Empty input {} — the shipped defaults2008.5 / 8.5 s
One whole voivodeship (opolskie), complete + reconciled61511.4 / 13.6 s
Three whole voivodeships in parallel, complete + reconciled1,9127.4 s (6 page requests, 0 retries)
Permits issued in the last 365 days20022.5 s
Masovia independents with an e-mail (3,060 rows read)61822.4 s
Permit-type filter, scanned client-side across the corpus2069.1 s — the slowest of 43 audited runs
The 17.5 MB official CSV export, fetched directmedian 4.4 s, range 3.8 – 21.4 s (17 runs)

About that range. The register's own file server is what varies, not this Actor: the identical 17.5 MB download measured 3.8 s on most runs and 21.4 s on two, and those two were runs we fired ten at a time at the same government host. Every one of 43 audited runs finished well inside Apify's 300-second health window, the slowest at 69.1 s.

The 1,912-row run is the completeness proof in miniature: lubuskie 578, opolskie 615, podlaskie 719, each stream's unique register ids exactly equal to the count the register declared, 0 duplicates, 0 dropped, register_complete: true, and 1,912 rows billed against 1,912 delivered.

Why there is no proxy by default

Same crawl, same build, three network settings:

Time
Direct out of the container22.4 s
Apify proxy, page size 20058.4 s
Apify proxy, page size 500336.9 s (~47 s per 1.8 MB page)

And on the 17.5 MB export: 4–5 s direct, 42.7 s proxied — with one proxied attempt timing out at 180 s and its retry pushing a 200-row run to 249 s. The register has no anti-bot of any kind and no rate limiting appeared across 100+ measured requests, so the proxy costs up to 15x and buys nothing. It is one click away if your own egress is ever blocked, and when it is on the Actor pins one proxy session per crawl stream — measured 20/20 successful with a 5.6 s p90, against 39/40 and a 47.3 s p90 unpinned. The log tells you, with these numbers, when the proxy is what is making a run slow. The bulk export is always fetched direct unless you explicitly ask otherwise, and that download is deadline-aware: if too little of the run's time budget is left it is skipped, the log says so, and it is not charged.

Charging limits are honoured exactly

Two runs with maxTotalChargeUsd set, verified against the platform's own counters:

CapExport joinDeliveredBilledActual spend
$0.10on32 rows32 rows + 1 join$0.1000
$0.05off (8 parallel streams)20 rows20 rows, 0 joins$0.0500
$0.03on4 rows4 rows + 1 join$0.0300
$0.02on0 rowsnothing at all$0.0000

Not one row over, and not one row delivered unbilled. Streams run concurrently, so the money budget is reserved synchronously before any batch can be pushed — reserving it after the fact let two streams each claim the same last 32 rows and hand over 64 for the price of 33, which is exactly what an earlier build did before this was measured and fixed. The one-time join charge is held back in whole rows, not in dollars, because (0.03 - 0.02) / 0.0025 is 3.9999999999999996 in binary floating point and a build that reserved in dollars quietly sold you 3 rows on a cap that pays for 4.

A cap that leaves no room for even one row (the $0.02 line above) delivers nothing and charges nothing — not even the join, which had already been downloaded — and the log says the charging limit did it and offers to turn the join off so the whole cap buys rows.

Memory

Peak RSS for a full 23,993-row crawl with the join on: 313 MB with the CSV export, 425 MB with the XML one — which is why the declared run size is 1024 MB. On CSV (the default), or with the join off, 512 MB is plenty and you can lower it in the run options.

The only failures observed in 100+ measured requests were socket aborts with no HTTP status and no rate-limit signature: transport, not throttling, and exactly what the retry-with-backoff is for. An HTTP 500 is never retried, because for this register a 500 means the query itself is malformed — the run fails loudly instead of hammering a government server.


🔒 Privacy — read this

This is a public state register and this Actor publishes only what the state publishes. It still contains personal data, and the defaults are set accordingly:

  • owners[].pesel — the Polish national identity number — is hard-dropped in code, under every setting, and can never appear in your dataset. It was null on all 23,993 rows and the state's own XSD excludes it from the public data scope. It is dropped from raw_json too.
  • The pharmacy manager block is OFF by default. manager_first_name, manager_last_name, manager_licence_number_pwz (the pharmacist's professional-practice number), manager_since and the deputy manager are only parsed when you switch includeManagerNames on. They are filled on 99.9% of active rows — and they are a named individual and their professional licence number.
  • owner_first_name / owner_last_name are natural persons. 35.8% of the register is owned by a sole trader, frequently registered at a home address. They are business-register entries, but they are still people.

You need your own lawful basis under GDPR to process any of that. Direct B2B marketing to a business's filed contact is usually defensible; building a profile of a named pharmacist is a different question. That call is yours, not ours.


Rejestr Aptek is an official public register published by the Centrum e-Zdrowia under Polish pharmaceutical law, with a documented XSD and a self-service bulk export the state offers for download. Access is not the moat here and this page will not pretend it is. Anyone can download the same 17.5 MB CSV. What this Actor sells is the work on top: 50 voivodeship spellings normalised to 16 and verified against the register's own partition, the active/dead split, the server-side filters that the register documents nowhere, the loss-free crawl order, a permit-date feed, and a row shape you can load straight into a CRM.

There is no robots.txt on either host — both answer /robots.txt with HTTP 200 and the Angular application shell, so there is no directive to honour or to argue about. The Actor is polite by default: 4 parallel streams, exponential backoff, and it never retries a request the server told it was malformed.

You are responsible for how you use the output, including GDPR and Polish direct-marketing rules (and, for anything touching medicines advertising, Prawo farmaceutyczne). Verify anything business-critical against the register itself — every row carries a source_url that returns that single record.


❓ FAQ

How many pharmacies are there? 23,993 records in the register on 2026-09-08; 13,733 are active. The rest are closed (9,950), suspended (166), temporarily closed (124) or pending (20).

How many actually have contact details? Of the 13,733 active: 98.4% have an e-mail, 97.8% a phone, 99.4% at least one of the two, and 99.8% carry the owner's NIP. Across the whole register including dead rows: 83.5% / 89.5% / 91.1% / 93.0%.

Does this cover independent pharmacies only, or chains too? Both, and it tells them apart. 7,225 distinct owner NIPs run the 13,733 active pharmacies; 2,120 of those owners have more than one and the largest has 101. 4,911 active pharmacies belong to an independent (sole trader, civil-law partnership or other non-registered person).

Can I get only newly licensed pharmacies? Yes — permitIssuedInLastDays, which the register honours as a real server-side date bound. Expect a small feed: 12 permits in the last 30 days, 46 in 90, 200 in 365. Six of the last 90 days' permits are still oczekująca — granted but not yet open.

Is the whole register really covered, or does paging drop rows? Really covered, and the Actor proves it per run. Sorted by register id, three independent full crawls each returned 23,993 unique ids, matching the official bulk export exactly. Every run reconciles the unique ids it saw against the count the register declared, and writes register_complete to RUN_SUMMARYfalse whenever a cap, a charging limit, a timeout or a failed page truncated the read, false on any run that delivered zero rows, and false when cross-run dedupe held matching rows back. register_complete_reason says which, in words.

That last pair is a bug we shipped and fixed. On build 0.1.8, a registerIds lookup of one pharmacy that happens to be nieaktywna, run with the default Active pharmacies only, read its single row to the end and reconciled perfectly (1 unique id against 1 declared, no cap, no timeout) — and then the status filter dropped that row. The record read delivered: 0 next to register_complete: true. A completeness flag is the one field a pipeline trusts blindly, so from build 0.1.9 it describes the delivered result and can never certify an empty dataset.

Re-verifying that fix on the live Actor turned up the same lie pointed the other way, fixed in build 0.1.11. Max pharmacies also acts as a stop signal, and it fires the moment the delivered count reaches it — including on the last row of a result set the crawl had already read to the end. So a run that delivered every matching row came back truncated: true, reason "the read was truncated (maxItems cap)", for a truncation that never happened (city: "Opole", maxItems: 113, 113 declared, 113 delivered, 0 dropped). The cap now counts against completeness only when it actually withheld a row, and a false "we cut this short" is treated as exactly as serious as a false "this is everything".

Do I need the official-export join? Only if you want the region hierarchy, TERYT codes, opening hours, launch date, coordinates or the manager. Everything else comes straight from the search endpoint. It costs one 17.5 MB download — measured over 17 Apify runs at a median of 4.4 s (range 3.8–21.4 s, the spread being the government file server, not this Actor), with a 200-row run finishing in 9.4 s — and $0.02 per run. It is on by default because a pharmacy row without a voivodeship is hard to use.

Which export file should I pick? CSV, unless you specifically need latitude/longitude — that is the XML file's one exclusive column and it is filled on just 21.9% of the register (34.5% of active rows). CSV is 17.5 MB against 54.1 MB, roughly six times faster, and carries more columns.

Are the coordinates real latitude and longitude? They are now. The state's XML calls its two coordinate attributes szerokoscGeograficzna and dlugoscGeograficzna — "geographic latitude" and "geographic longitude" — and they are neither: they hold PUWG-1992 (EPSG:2180) grid metres, and the two are swapped against their own names. Scored against each row's own voivodeship centroid over the 1,770 coordinate-bearing rows in the first 30 MB of the file, reading dlugoscGeograficzna as the northing and szerokoscGeograficzna as the easting puts 97.7% of pharmacies within 120 km of the right region (median 49.6 km); the name-order reading manages 27.7% (median 182.5 km). This Actor does the inverse transverse-Mercator itself and emits WGS84 degrees in latitude / longitude, keeps the raw grid pair in puwg92_x / puwg92_y so you can check the conversion, and drops the pair entirely — rather than plotting it somewhere wrong — when the result falls outside Poland (10 rows of 5,268, which is why the fill rate is 21.9% and not 22.0%). Spot check: register id 1232425, Bydgoszcz 85-791, converts to 53.167174 N / 18.170671 E.

Why is website almost empty? Because the register almost never has one — 1.6% of rows, 2.1% of active ones. We publish the real number rather than implying a web-presence dataset.

What is TERC / SIMC / ULIC? Poland's official TERYT statistical codes for the commune, the locality and the street. They are the clean key for joining to GUS statistics and to most Polish geodata. Filled on ~50% of rows and only available via the export join.

Can I monitor the register for changes? Yes: changedInLastDays plus dedupeAcrossRuns: true, on a schedule. The seen-id set lives in a named key-value store so it survives between runs. Volumes: 38 records changed in the last day, 381 in 7 days, 652 in 30, 1,862 in 90.

Is the data in Polish? The register's own values are, and they are preserved exactly. English is added alongside for the three closed vocabularies that matter: status_en, pharmacy_type_en and owner_legal_form_en.

Does it need a proxy? No, and it is off by default. The register has no anti-bot of any kind, and the proxy measured up to 15x slower on the megabyte-sized pages this register returns: the same Masovia crawl took 22.4 s direct and 336.9 s proxied. Switch it on if your own egress is blocked — the Actor then pins one session per stream, which measured 20/20 with a 5.6 s p90 — and the log tells you when the proxy is what is making a run slow.

What happens if my charge limit runs out mid-run? Delivery stops on the row that would have exceeded it. Both budgets — your row cap and your money — are reserved before any batch is written, so you are never billed for rows you did not receive and never handed rows you were not billed for. Verified: a $0.10 cap delivered and billed exactly 32 rows plus the one join, and a $0.05 cap with the join off delivered and billed exactly 20. The log names the charging limit as the cause instead of blaming the register.