# Changelog of Pracuj Scraper \[💰$2] — Poland's #1 Job Board (`blackfalcondata/pracuj-scraper`) Actor

- **URL**: https://apify.com/blackfalcondata/pracuj-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/blackfalcondata/pracuj-scraper.md

## Changelog

### 0.1.x — 2026-10-06

- Fixed: a run with a maximum cost per run could end as aborted. The run kept
  working past what the maximum cost pays for, and the platform then stopped it
  midway. The run now works out up front how many results the maximum cost
  covers, says so in the log, delivers exactly those and finishes normally.
- Fixed: a long run that the platform moved to another machine fetched full
  details again for listings it had already delivered. It now continues straight
  from where it left off.
- Improved: large runs use less memory, so a run of several thousand listings
  with full details fits comfortably in the smallest memory setting.
- Fixed: a run given a single job-offer link could fail on a brief hiccup while
  fetching that offer. The fetch is now retried a few times before the run
  gives up.

### 0.1.72 — 2026-09-28

- Fixed: an English city name such as `Warsaw` found no jobs. It is now searched as `Warszawa` (also Cracow, Breslau, Danzig, Posen, Stettin).
- Fixed: `remote` in the Location field found no jobs. It now searches remote jobs, the same as Work Mode: Remote.
- Improved: a run that finds nothing now says why — including when several keywords were typed into one Search Term, which the site searches as one text.

### 0.1.x — 2026-08-05 (b)

- Fixed: large runs could deliver nothing at all. Results were built up in memory
  and written only at the very end, so when the platform moved a long run to
  another machine — which it does routinely — everything collected up to that
  point was lost and the run reported success with no results. Results are now
  written continuously as they are scraped, and a run that is moved picks up
  where it left off instead of re-delivering (and re-charging) what it already
  sent. Measured before the fix: 34 minutes, 10,926 listings collected, 0
  delivered.
- Fixed: a run that could not reach the source at all now fails instead of
  reporting success with zero results. A partial scrape still succeeds.
- Changed: the location-radius filter is applied before Max Results rather than
  after, so a radius search returns the number of jobs asked for instead of
  however many of the first N happened to fall inside it.

### 0.1.x — 2026-08-05

- Added: `categories` input — narrow a search to one or more of pracuj.pl's own
  job categories. Picking "Sales" together with "Customer service" returns the
  same listings as the sales sub-portal, sprzedaz.pracuj.pl; "IT - Software
  development" plus "IT - Administration" matches it.pracuj.pl. Applied at the
  source, so it costs no extra requests and combines with the search term and the
  other filters.
- Added: a category on its own is now enough to run. No search term is needed to
  scrape a whole category, which is how the sub-portals are browsed.
- Fixed: pasting a sub-portal URL (sprzedaz.pracuj.pl, it.pracuj.pl) scraped the
  entire portal instead of that category. Those hosts carry their category in the
  hostname and nowhere in the URL, so the filter was invisible to the parser and
  the run spent its budget on unrelated offers.
- Changed: the memory ceiling is no longer pinned at the default, so a large run
  can be given more memory instead of being clamped back to 256 MB.

### 0.1.x — 2026-07-25 (b)

Follow-up to the release below, from an adversarial audit of it.

- Fixed: a weekly salary was reported as a daily one. The time-unit match tested
  for "day" before "week", and the Polish word for week contains the day stem.
- Fixed: when an offer prices each contract type differently, the structured
  salary is no longer taken from whichever contract happened to be listed first —
  which contradicted `salaryText` and flipped with array order. The aggregate
  range now stands unless the offer states a single salary.
- Fixed: pasting several search URLs only scraped the first one. A page is
  fetched whole, so the first URL consumed the entire result budget; each search
  now gets a share of it.
- Fixed: a search term and a pasted search URL are now both searched. The term
  used to be dropped whenever any URL was pasted.
- Fixed: a pasted URL whose filters could not be read is skipped instead of being
  treated as an unfiltered search of the whole portal.
- Fixed: `startUrls` no longer requires a search term to be set as well.
- Fixed: with `emitExpired`, a listing we could not reach because of a network or
  server error is no longer reported as expired — only a confirmed removal is.
- Fixed: expired markers now count towards Max Results instead of being emitted
  without limit, and carry their own `expiredAt` field rather than overwriting
  `expirationDate`, which holds the expiry the source advertised.
- Fixed: pasted offer URLs beyond Max Results are no longer fetched.
- Changed: detail requests no longer retry twice over (the HTTP layer already
  retries), so a struggling source is not hit harder than necessary.
- Changed: run logs no longer include internal retrieval details.
- Rejected: URLs that are not http(s), and offer ids too large to represent
  exactly, which previously resolved to a different listing.

### 0.1.x — 2026-07-25

- Added: `employerType` input — return only offers posted directly by the hiring
  company, or only offers posted by recruitment agencies. Applied at the source,
  so it costs no extra requests and works without `includeDetails`.
- Added: `isDirectlyFromEmployer` output field. Resolved from the offer detail
  when details are fetched, and implied for every record of a run narrowed with
  `employerType`; `null` when neither applies.
- Added: `startUrls` input — paste pracuj.pl offer URLs (fetched individually)
  or search/listing URLs (paginated like a search). Filters in a pasted URL take
  precedence; the input filters still apply to whatever the URL leaves unset.
- Added: `emitExpired` now emits rows for listings that have disappeared since a
  previous incremental run (`jobId`, `changeType: "EXPIRED"`, seen-timestamps).
- Fixed: structured salary, region, coordinates, contract types, work modes,
  work schedules and position levels were read from an `attributes.employmentTypes`
  path that does not exist on the detail response, so detail enrichment silently
  never filled them. `salaryMin`/`salaryMax`/`salaryCurrency`/`salaryPeriod` now
  come from the offer's actual salary data instead of only a parsed display string.
- Changed: detail fetches now carry a per-request deadline, so one stalled
  request can no longer hold a run open until the platform time limit.

### 0.1.x — 2026-04-14

- Added: `descriptionHtml`, `descriptionMarkdown` output fields (triple-format descriptions for RAG/LLM pipelines)
- Added: `contentHash` output field for stable change detection

### 0.1.x — 2026-04-14

- Added: cross-run repost detection (`isRepost`, `repostOfId`, `repostDetectedAt`)
- Added: `skipReposts` input to exclude detected reposts from output

### \[0.1.0] - 2026-03-22

#### Initial release

- Search Polish job listings on pracuj.pl by keyword, location, and filters
- Salary extraction (min, max, currency, period)
- Detail enrichment: full description, responsibilities, requirements, benefits, technologies
- Contact info: apply URL and phone when available
- Incremental mode: only new or changed listings on subsequent runs
- Compact output mode for AI-agent and MCP workflows
- Filters: contract type, work mode, position level
