# Pracuj.pl Job Scraper — Poland Jobs, Salaries & Companies (`corvuslab/pracuj-scraper`) Actor

Scrape pracuj.pl, Poland's #1 job board. Export structured job offers with salary (min/max/currency), contract type, work mode, tech stack, requirements, benefits, company data, apply links and AI summaries. Incremental monitoring + notifications built in.

- **URL**: https://apify.com/corvuslab/pracuj-scraper.md
- **Developed by:** [Corvuslab](https://apify.com/corvuslab) (community)
- **Categories:** Jobs, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.90 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Pracuj.pl Job Scraper — Poland Jobs, Salaries & Company Data

> **Turn every live pracuj.pl job offer into clean, structured data — in seconds, at scale.**

Scrape job offers from **[pracuj.pl](https://www.pracuj.pl/)**, Poland's #1 job board, and get
back everything that matters: **structured salary (min / max / currency)**, contract type, work
mode, tech stack, requirements, benefits, company details and the direct apply link. Search by
keyword and filters or paste your own pracuj.pl URL — no code required. Export to **JSON, CSV,
Excel or the API**, or pipe it straight into Google Sheets, n8n, Make or an AI agent (MCP).

**Why this scraper**

- ⚡ **Fast & low-cost** — pure HTTP, no headless browser, so large runs stay cheap.
- 🧾 **Rich, typed records** — 30+ structured fields per offer, not raw HTML.
- ♻️ **Cheap to monitor** — incremental mode re-scrapes only what changed (see below).
- 🔔 **Notifications built in** — Telegram, Slack, Discord, WhatsApp or any webhook.
- 🤖 **AI- & API-ready** — compact output, MCP-friendly, one-click integrations.

***

### ✨ Key features

- 🇵🇱 **The whole of pracuj.pl** — search across every category and all 16 voivodeships, or paste
  a ready-made pracuj.pl search URL and let the scraper page through it.
- 💰 **Structured salary** — not just the raw text: `salaryFrom`, `salaryTo`, `salaryCurrency`,
  time unit and gross/net kind, broken out for instant benchmarking.
- 🧠 **Full job details** — description (text / HTML / Markdown), responsibilities, required &
  optional requirements, benefits, tech stack, required languages, IT specializations and the
  employer's own **AI summary** of the role.
- 🏢 **Company & apply data** — company name, ID, profile URL, logo, direct-employer flag, the
  **direct apply link**, reference number and one-click-apply flag.
- 🎚️ **Rich filters** — keyword, region, category, seniority, contract type (B2B, employment,
  mandate…), work mode (remote / hybrid / on-site), work schedule, "posted within" and salary-only.
- ♻️ **Incremental monitoring** — schedule a run and get only what changed (NEW / UPDATED /
  EXPIRED), with reposted-offer detection; unchanged offers are skipped before their page is fetched.
- 🔔 **Notifications** — push new offers to Telegram, Slack, Discord, WhatsApp or any webhook
  (n8n / Make / Zapier).
- 🤖 **AI-ready** — compact mode + drop-empty-fields mode keep payloads lean for LLMs and MCP.
- ⚡ **Fast & cheap** — pure HTTP, no headless browser.

***

### 📤 Example output

```json
{
  "id": "1004957230",
  "title": "Senior Software Engineer (Python, GenAI & Agentic AI)",
  "url": "https://www.pracuj.pl/praca/senior-software-engineer-python-genai-agentic-ai-warszawa,oferta,1004957230",
  "company": "IN4GE sp. z o.o.",
  "companyProfileUrl": "https://pracodawcy.pracuj.pl/company/1074095560",
  "location": "Warszawa",
  "city": "Warszawa",
  "region": "mazowieckie",
  "country": "Polska",
  "isRemote": true,
  "salaryText": "140–170 zł netto (+ VAT) / godz.",
  "salaryFrom": 140.0,
  "salaryTo": 170.0,
  "salaryCurrency": "PLN",
  "salaryTimeUnit": "godzinowo",
  "salaryKind": "netto (+ VAT)",
  "contractTypes": ["Kontrakt B2B"],
  "workModes": ["Praca zdalna"],
  "positionLevels": ["Starszy specjalista / Starsza specjalistka (senior)"],
  "technologies": ["Python"],
  "technologiesOptional": ["LangGraph", "LangChain", "Agentic AI", ".NET"],
  "requiredLanguages": ["angielski"],
  "requirements": ["Minimum 5 lat doświadczenia komercyjnego w obszarze Software Engineering.", "Bardzo dobra znajomość Python."],
  "responsibilities": ["Projektowanie, rozwój i utrzymanie rozwiązań opartych o Generative AI, Agentic AI oraz RAG."],
  "offered": ["możliwość pracy zdalnej", "elastyczny czas pracy"],
  "categories": ["Programowanie", "Architektura"],
  "aiSummary": "Masz minimum 5 lat doświadczenia w software engineering oraz praktyczną znajomość Python...",
  "applyUrl": "https://www.pracuj.pl/aplikuj/...,oferta,1004957230",
  "postedDate": "2026-08-05T15:00:00Z",
  "expirationDate": "2026-08-08T21:59:59Z",
  "source": "pracuj.pl"
}
```

Descriptions are also available as **HTML** and **Markdown**, not just plain text.

***

### 📚 What data can you extract?

- **Core** — id, title, URL, company, location / city / region / country, remote flag, categories,
  posted & expiration dates.
- **Salary** — raw text plus structured `salaryFrom` / `salaryTo` / `salaryCurrency`, time unit
  and gross-vs-net kind.
- **Details** (with *Fetch full details*) — description (text / HTML / Markdown), responsibilities,
  required & optional requirements, benefits offered, tech stack (required + optional), required
  languages, position levels, contract types, work modes and the employer's AI summary.
- **Company & apply** — company ID, profile URL, logo, direct-employer flag, direct apply link,
  reference number and one-click-apply flag.

Every field is present in standard mode (missing values are `null`); **compact mode** returns the
core fields only, for lean AI/MCP payloads.

***

### ⚙️ Input

Configure it in the visual editor — no code needed — or pass JSON via the API.

| Field | What it does |
|---|---|
| `query` | One or more keywords; each runs as its own search, merged & de-duplicated. |
| `startUrls` | Paste ready-made pracuj.pl search URLs or individual offer URLs. |
| `region` / `category` / `positionLevel` / `contractType` / `workMode` / `workSchedule` | Multi-select filters applied on the site. |
| `postedWithin` / `withSalary` | Narrow to fresh or salaried offers only. |
| `includeDetails` | Off = fast listing-only; on = full description, requirements, benefits, structured salary, apply link. |
| `incrementalMode` | Emit only what changed since the previous run. |
| `maxResults` | Cap records per run (0 = unlimited). |
| `proxyConfiguration` | Optional — works without a proxy at typical volumes. |

…and **31 inputs** in total — the table shows the essentials; the rest cover description formatting
(`descriptionFormat`, `descriptionMaxLength`), AI/compact output (`compact`, `excludeEmptyFields`),
notification channels (Telegram / Slack / Discord / webhook) and advanced tuning, all in the visual
editor.

**Example inputs**

```json
{ "query": "python developer", "region": ["mazowieckie"], "maxResults": 100 }
```

```json
{ "query": "data engineer", "workMode": ["Praca zdalna"], "withSalary": true, "includeDetails": true }
```

```json
{ "query": "devops", "incrementalMode": true, "stateKey": "devops-pl", "telegramToken": "...", "telegramChatId": "..." }
```

***

### 💡 Use cases

- **Recruitment & sourcing** — build a live feed of open roles and companies hiring in Poland,
  filtered to your niche.
- **Salary benchmarking** — aggregate `salaryFrom` / `salaryTo` by role, seniority, region,
  contract type or tech stack.
- **Lead generation** — find companies actively hiring (a strong buying signal) with their profile
  URLs and apply channels.
- **Market & talent research** — track demand for skills, technologies and languages across the
  Polish market over time.
- **Job boards & aggregators** — enrich your own listings with structured pracuj.pl data.

***

### ♻️ Incremental monitoring — pay for changes, not repeats

Schedule the actor and turn on **incremental mode**: each run compares against the last and emits
only **NEW / UPDATED / EXPIRED** offers — unchanged offers are skipped *before* their page is even
fetched, so a daily watch costs a fraction of a full re-scrape.

| Daily churn | of 1,000 tracked | billable offers | you save |
|---|---|---|---|
| 5 % | 1,000 | 50 | **95 %** |
| 15 % | 1,000 | 150 | **85 %** |
| 30 % | 1,000 | 300 | **70 %** |

The first run seeds the baseline and bills in full; every run after that bills only the delta.

***

### 🚀 How to run it

1. Enter a **search keyword** (e.g. `python developer`) and/or pick filters — or paste a pracuj.pl URL.
2. Set **Max results** and choose whether to **fetch full details**.
3. (Optional) Turn on **incremental mode** and a **notification** channel, then **Schedule** it.
4. Click **Start**.
5. Download the data as **JSON, CSV or Excel**, or pull it from the **API**.

New to Apify? Create a free account — it comes with monthly credit, no credit card required.

***

### 🔌 Integrations & export

Export to **JSON, CSV, Excel** or an HTML table, or pull from the **REST API** and the
**JavaScript / Python** clients. Runs on a **schedule**, connects to **Google Sheets, Slack, Make,
Zapier and n8n**, and works as an **MCP tool** for AI agents — compact mode plus
`descriptionMaxLength` keeps token usage small.

***

### ❓ FAQ

**Do I need to log in or provide a proxy?** No — it works out of the box with no proxy. Apify Proxy
is available under Advanced if you ever want it for very high-volume runs.

**Can I get only new offers on a schedule?** Yes — turn on **incremental mode** and schedule the
actor. Each run emits only NEW / UPDATED (and optionally EXPIRED) offers, and can notify your
Telegram / Slack / Discord / WhatsApp / webhook.

**What formats can I export?** JSON, CSV, Excel, HTML table, or via the API.

**Is the description available as HTML / Markdown?** Yes — choose plain text, HTML, Markdown, or
all three, with an optional length cap.

**How many offers can I get?** As many as the search returns — set `maxResults` (0 = unlimited).
Leave the keyword and location empty to sweep the whole board.

**Is scraping this legal?** The scraper collects only **publicly available** job-posting data. You
are responsible for how you use it, including compliance with the GDPR and pracuj.pl's terms.

***

### ⚖️ Disclaimer

This actor accesses only publicly available data on pracuj.pl. You are responsible for how you use
the extracted data — in particular any personal information — and for complying with the site's
terms and applicable law (including the GDPR where it applies). Not affiliated with, endorsed by,
or sponsored by Grupa Pracuj S.A.

***

**Keywords:** pracuj scraper · pracuj.pl scraper · pracuj api · poland jobs scraper · polish job board · praca scraper · oferty pracy scraper · IT jobs poland · salary data poland · recruitment sourcing · lead generation · labour market research · job monitoring · export to CSV/Excel/JSON · no-code job scraper · MCP tool for AI agents

# Actor input Schema

## `query` (type: `string`):

Keywords to search for (e.g. "programista", "marketing manager"). Separate multiple searches with commas — each runs as its own search and the results are merged and de-duplicated.

## `startUrls` (type: `array`):

Paste pracuj.pl search/listing URLs (with your own filters applied on the site) or individual offer URLs. These are scraped in addition to the keyword search.

## `maxResults` (type: `integer`):

Maximum number of job offers to return. Set 0 for unlimited (bounded by how many the search actually has).

## `ignoreUrlFailures` (type: `boolean`):

Skip start URLs that cannot be interpreted instead of failing the whole run.

## `region` (type: `array`):

Polish voivodeship(s) to limit the search to.

## `category` (type: `array`):

Top-level pracuj.pl job categories.

## `positionLevel` (type: `array`):

Seniority / position level.

## `contractType` (type: `array`):

Type of employment contract.

## `workMode` (type: `array`):

On-site, hybrid, remote or mobile.

## `workSchedule` (type: `array`):

Full-time, part-time or additional/temporary.

## `postedWithin` (type: `string`):

Only offers published within this window.

## `withSalary` (type: `boolean`):

Return only offers that disclose a salary.

## `includeDetails` (type: `boolean`):

Fetch each offer's page for the richer fields (description, requirements, benefits, technologies, structured salary, apply URL). Turn off for the fastest, cheapest listing-only runs.

## `descriptionFormat` (type: `string`):

Which representation(s) of the job description to include.

## `descriptionMaxLength` (type: `integer`):

Truncate each description variant to at most this many characters. Leave empty for full text.

## `compact` (type: `boolean`):

Emit only the core fields (title, company, salary, location, tech, dates). Ideal for AI agents and MCP clients.

## `excludeEmptyFields` (type: `boolean`):

Remove null, empty-string and empty-array fields from each record.

## `incrementalMode` (type: `boolean`):

Track state between runs and tag every offer with a changeType (NEW / UPDATED / UNCHANGED / EXPIRED). Unchanged offers are skipped before their detail page is fetched, so recurring runs stay cheap.

## `stateKey` (type: `string`):

Stable name for the tracked search. Leave empty to derive one automatically from your search settings.

## `emitUnchanged` (type: `boolean`):

Also emit offers that have not changed since the previous run.

## `emitExpired` (type: `boolean`):

Emit offers that were present last run but are gone now (changeType = EXPIRED).

## `telegramToken` (type: `string`):

Bot token from @BotFather.

## `telegramChatId` (type: `string`):

Chat or channel ID, e.g. "-100123456789" or "@yourchannel".

## `slackWebhookUrl` (type: `string`):

Slack incoming-webhook URL.

## `discordWebhookUrl` (type: `string`):

Discord incoming-webhook URL.

## `webhookUrl` (type: `string`):

Any HTTPS endpoint. Receives a JSON POST with the matched offers — works with n8n, Make and Zapier.

## `webhookHeaders` (type: `object`):

Extra headers for the webhook request, e.g. {"Authorization": "Bearer xyz"}.

## `notificationLimit` (type: `integer`):

How many offers to include in each notification message.

## `proxyConfiguration` (type: `object`):

pracuj.pl blocks datacenter/cloud IPs, so a Polish residential proxy is used by default and is recommended. The scraper automatically rotates to a fresh IP if one is refused.

## `detailConcurrency` (type: `integer`):

How many offer detail pages to fetch in parallel.

## `maxRequestRetries` (type: `integer`):

How many times to retry a failed request before giving up on it.

## Actor input object example

```json
{
  "query": "programista, data engineer",
  "maxResults": 25,
  "ignoreUrlFailures": true,
  "postedWithin": "",
  "withSalary": false,
  "includeDetails": true,
  "descriptionFormat": "all",
  "compact": false,
  "excludeEmptyFields": false,
  "incrementalMode": false,
  "emitUnchanged": false,
  "emitExpired": false,
  "notificationLimit": 5,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "PL"
  },
  "detailConcurrency": 8,
  "maxRequestRetries": 3
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `allItems` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "programista"
};

// Run the Actor and wait for it to finish
const run = await client.actor("corvuslab/pracuj-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "query": "programista" }

# Run the Actor and wait for it to finish
run = client.actor("corvuslab/pracuj-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "programista"
}' |
apify call corvuslab/pracuj-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=corvuslab/pracuj-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/S4d9snTQ6PsgsBz6C/builds/DfhBLrrPUFbNzVZic/openapi.json
