# Habr Career Scraper: Вакансии Хабр Карьера (`getascraper/habr-career-scraper`) Actor

Scrape Habr Career (Хабр Карьера) tech job vacancies with structured skill tags, specialization, and full description, plus real, large-sample Russian IT salary benchmarks by seniority level from Habr's own public salary survey. Includes a new-postings monitor for a saved search. No login required.

- **URL**: https://apify.com/getascraper/habr-career-scraper.md
- **Developed by:** [GetAScraper](https://apify.com/getascraper) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.05 / 1,000 job postings

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## 💼 Habr Career Scraper: Вакансии Хабр Карьера

<table width="100%">
<tr>
<td style="padding:24px 28px;background:#F3FAEE;border:1px solid #CBE8B4;border-top:4px solid #4C9A2A;border-radius:12px">
<span style="font-size:23px;font-weight:800;color:#1C1917;line-height:1.3">Real Russian IT salary data, not just job postings.</span><br>
<span style="font-size:15px;color:#57534E;line-height:1.6">Scrape tech vacancies from Хабр Карьера (Habr Career) with structured skill tags and specialization, then cross-check them against Habr's own public salary survey: real pay by seniority level, drawn from a 41,000-plus data point sample.</span>
</td>
</tr>
</table>

<table width="100%">
<tr>
<td style="padding:14px 12px;width:25%;background:#FFFFFF;border:1px solid #CBE8B4;border-radius:10px 0 0 10px;vertical-align:top">
<span style="font-size:15px;font-weight:800;color:#2F6B17">💰 Real salary benchmarks</span><br>
<span style="font-size:12px;color:#57534E">Min, p25, median, p75, max, and sample size per seniority level, straight from Habr's own survey.</span>
</td>
<td style="padding:14px 12px;width:25%;background:#FFFFFF;border:1px solid #CBE8B4;border-left:none;vertical-align:top">
<span style="font-size:15px;font-weight:800;color:#2F6B17">🏷️ Structured skill tags</span><br>
<span style="font-size:12px;color:#57534E">Each vacancy's skills come back as named tags, not a flat keyword string.</span>
</td>
<td style="padding:14px 12px;width:25%;background:#FFFFFF;border:1px solid #CBE8B4;border-left:none;vertical-align:top">
<span style="font-size:15px;font-weight:800;color:#2F6B17">🎯 Specialization matching</span><br>
<span style="font-size:12px;color:#57534E">A vacancy's division lines up directly with the salary survey's own categories.</span>
</td>
<td style="padding:14px 12px;width:25%;background:#FFFFFF;border:1px solid #CBE8B4;border-left:none;border-radius:0 10px 10px 0;vertical-align:top">
<span style="font-size:15px;font-weight:800;color:#2F6B17">🔔 New-posting alerts</span><br>
<span style="font-size:12px;color:#57534E">Run it on a schedule and get only the vacancies you haven't seen before.</span>
</td>
</tr>
</table>

***

### 🔍 What does Habr Career Scraper do?

Habr Career (Хабр Карьера) is Russia's leading job board for tech roles: developers, QA,
analysts, DevOps, and IT management. This actor searches its public listings and returns clean,
structured data for every matching vacancy, no account or login required.

Each vacancy comes back with title, employer name, logo and rating, salary when disclosed,
location, remote flag, employment type, seniority level, full description, posting date, and
expiry date. Skills and specialization arrive as structured tags with canonical names, not a
loose keyword string.

The part no other Russian job scraper offers: turn on salary benchmarks and get a second dataset
straight from Habr's own public compensation survey. Real minimum, 25th percentile, median, 75th
percentile, and maximum pay by seniority level (Junior through Lead), each backed by a real
sample size in the tens of thousands. Because Habr uses the same specialization categories for
both its job listings and its salary survey, a scraped vacancy's division lines up with a
benchmark row with no extra matching work.

***

### 👥 Who uses it?

**Tech recruiters tired of HH.ru's noise** - "HH.ru mixes every industry together. We only place
developers and DevOps engineers, so we run this against Habr Career instead and get an IT-only
feed with skill tags we can filter on directly."

**HR and comp teams benchmarking offers** - "Before we send an offer, we check the Salary
Benchmarks dataset for that seniority level. It tells us in seconds whether our number is
competitive or we're about to lose a candidate."

**Candidates sanity-checking a job offer** - "I pulled the Middle-level salary range before my
final interview. Knowing the real median gave me a number to negotiate from instead of guessing."

**Existing HH.ru Jobs Scraper users** - "We already run [HH.ru Jobs
Scraper](https://apify.com/getascraper/hh-ru-jobs-scraper) for general hiring. This actor is the
tech-specific complement: same shape of data, plus salary benchmarks HH.ru doesn't have."

***

### 🚀 How to use it

<table width="100%">
<tr>
<td style="padding:16px 14px;width:33%;background:#F3FAEE;border:1px solid #CBE8B4;border-radius:10px 0 0 10px;vertical-align:top">
<span style="font-size:12px;font-weight:800;color:#4C9A2A;letter-spacing:1px">STEP 1</span><br>
<span style="font-size:14px;font-weight:700;color:#1C1917">Set your search</span><br>
<span style="font-size:12px;color:#57534E">Enter keywords, a skill, a specialization, or a city. Salary benchmarks stay on by default.</span>
</td>
<td style="padding:16px 14px;width:33%;background:#F3FAEE;border:1px solid #CBE8B4;border-left:none;vertical-align:top">
<span style="font-size:12px;font-weight:800;color:#4C9A2A;letter-spacing:1px">STEP 2</span><br>
<span style="font-size:14px;font-weight:700;color:#1C1917">Run in the cloud</span><br>
<span style="font-size:12px;color:#57534E">The actor reads Habr Career's public pages directly. No login, no API key needed.</span>
</td>
<td style="padding:16px 14px;width:33%;background:#F3FAEE;border:1px solid #CBE8B4;border-left:none;border-radius:0 10px 10px 0;vertical-align:top">
<span style="font-size:12px;font-weight:800;color:#4C9A2A;letter-spacing:1px">STEP 3</span><br>
<span style="font-size:14px;font-weight:700;color:#1C1917">Download your data</span><br>
<span style="font-size:12px;color:#57534E">Export vacancies and salary benchmarks as JSON, CSV, or Excel, or connect to Sheets and BigQuery.</span>
</td>
</tr>
</table>

Want a standing watch instead of a one-off pull? Turn on **Alert on new listings** and give it a
State Name. Every scheduled run after that returns only vacancies posted since the last check.

***

### ⚙️ Input

| Field | Type | Required | Description |
|---|---|---|---|
| `keywords` | string | No | Free-text search query (e.g. `python`, `backend`). |
| `vacancyUrls` | array of URLs | No | Specific vacancy URLs or numeric IDs to scrape directly instead of searching. |
| `division` | enum | No | Specialization to browse, e.g. `development/backend`. Takes priority over Skills and Keywords/City. |
| `skills` | array | No | A single skill tag to filter by, e.g. `docker`, `kubernetes`. Only the first value is used. |
| `city` | enum | No | Russian city to filter by. Combines with Keywords. |
| `proxyConfiguration` | proxy | No | Apify proxy settings. Datacenter proxy is the default and sufficient. |
| `includeSalaryBenchmarks` | boolean | No | Fetch Habr's own salary survey once per run and add it as a second dataset. Default: on. |
| `maxItems` | integer | No | Maximum number of vacancies to return. Default: `20`. |
| `onlyNewListings` | boolean | No | Return only vacancies not seen in a previous run with the same State Name. Default: off. |
| `stateName` | string | No | Identifies which saved search the new-listings check belongs to. Default: `default`. |
| `resetState` | boolean | No | Clears saved new-listings history for State Name before this run. Default: off. |

***

### 📤 Data table

This actor produces two datasets: **Vacancies** (the default output) and **Salary Benchmarks**
(when Include Salary Benchmarks is on).

#### Vacancies

| Field | Type | Description |
|---|---|---|
| `vacancyId` | string | Unique Habr Career vacancy ID. |
| `title` | string | Vacancy title. |
| `url` | string | Link to the vacancy page. |
| `salaryFrom` | number | Minimum disclosed salary. Omitted when not disclosed. |
| `salaryTo` | number | Maximum disclosed salary. Omitted when not disclosed. |
| `salaryCurrency` | string | Salary currency code. |
| `employerName` | string | Employer company name. |
| `employerLogoUrl` | string | URL of the employer's logo image. |
| `employerRating` | number | Employer star rating from Habr's own company reviews. |
| `isTrustedEmployer` | boolean | True when Habr marks the employer as accredited. |
| `address` | string | Vacancy location. |
| `isRemote` | boolean | True when the vacancy is remote-friendly. |
| `employmentType` | string | Employment type as shown on Habr (e.g. full-time). |
| `qualification` | string | Seniority level (e.g. Junior, Middle, Senior, Lead). Matches the Salary Benchmarks `level` field. |
| `skills` | array | Structured skill tags, each with a name and canonical slug. |
| `division` | object | Primary specialization, with a name and slug matching the Salary Benchmarks taxonomy. |
| `description` | string | Full vacancy description. |
| `datePosted` | string | Date the vacancy was first published. |
| `validThrough` | string | Date the vacancy listing expires. |
| `scrapedAt` | string | ISO 8601 timestamp of when the row was scraped. |

#### Salary Benchmarks

| Field | Type | Description |
|---|---|---|
| `level` | string | Seniority level: All, Intern, Junior, Middle, Senior, or Lead. |
| `title` | string | Habr's own description of this breakdown. |
| `min` | number | Minimum monthly pay in the sample. |
| `p25` | number | 25th percentile monthly pay. |
| `median` | number | Median monthly pay. |
| `p75` | number | 75th percentile monthly pay. |
| `max` | number | Maximum monthly pay in the sample. |
| `sampleSize` | number | Real number of survey responses this row is computed from. |
| `currency` | string | Always RUB. |
| `asOfPeriod` | string | The most recent period Habr's survey reports, so you know how fresh the numbers are. |
| `scrapedAt` | string | ISO 8601 timestamp of when this snapshot was taken. |

***

### 💰 Pricing

This actor charges per vacancy returned. Salary benchmark rows are included at no extra charge.
Empty runs cost nothing. There are no monthly subscriptions or seat fees.

***

### ⭐ Enjoying Habr Career Scraper?

<table width="100%">
<tr>
<td style="padding:20px 24px 14px;background:#F3FAEE;border:1px solid #CBE8B4;border-left:5px solid #4C9A2A;border-radius:10px 10px 0 0">
<span style="font-size:20px;letter-spacing:4px">⭐ ⭐ ⭐ ⭐ ⭐</span><br>
<span style="font-size:17px;font-weight:800;color:#1C1917">If the salary benchmarks helped you make a real hiring or negotiation decision, we'd love to hear it.</span><br>
<span style="font-size:14px;color:#57534E">A 5-star rating takes 10 seconds and helps other recruiters and HR teams find it. Your feedback also tells us what to build next.</span>
</td>
</tr>
<tr>
<td style="padding:0;background:#4C9A2A;border:1px solid #CBE8B4;border-top:none;border-radius:0 0 10px 10px;text-align:center">
<a href="https://apify.com/getascraper/habr-career-scraper/reviews" style="display:block;padding:13px 16px;color:#FFFFFF;text-decoration:none;font-weight:800;font-size:15px;letter-spacing:0.3px">★&nbsp;&nbsp;Rate this Actor on Apify</a>
</td>
</tr>
</table>

***

### ❓ FAQ

**Does this actor require a Habr Career account or API key?**
No. It reads the public Habr Career website directly. No registration and no API key are needed
to run it.

**Как узнать реальную зарплату разработчика на Хабр Карьере?** (How do I find a developer's real
salary on Habr Career?)
Turn on Include Salary Benchmarks in the input form. The actor adds a second dataset with real
minimum, median, and maximum pay for every seniority level, sourced from Habr's own public
survey, no guessing required.

**Why do some vacancies show no salary fields?**
Habr lets employers hide their salary offer. When it isn't disclosed, salary fields are left out
of that row entirely rather than filled with a guess or a zero.

**Can I combine a keyword search with a specialization filter?**
Habr Career's own site only applies one search filter at a time. If you set Division, it takes
priority; otherwise Skills takes priority over Keywords and City. The actor logs which filters it
used so you always know what ran.

***

### 🔗 Other actors

- [HH.ru Jobs Scraper: Парсер вакансий HH.ru](https://apify.com/getascraper/hh-ru-jobs-scraper) ↗ - Search Russia's largest general job board by keyword, city, salary, and experience level.
- [Auto.ru Scraper: Парсер Авто.ру](https://apify.com/getascraper/autoru-scraper) ↗ - Scrape used-car listings from Russia's largest car marketplace with price-drop alerts.
- [Avito Auto Scraper: Парсер Авито Авто](https://apify.com/getascraper/avito-auto-scraper) ↗ - Extract used-car listings from Avito, Russia's largest classifieds site.
- [Yandex Realty Scraper: Яндекс Недвижимость](https://apify.com/getascraper/yandex-realty-scraper) ↗ - Extract property listings and pricing from Yandex Real Estate across Russia.
- [NoFluffJobs scraper: tech jobs & salaries](https://apify.com/getascraper/nofluffjobs-scraper) ↗ - Scrape tech job listings and transparent salary ranges from Poland's leading IT job board.

# Actor input Schema

## `keywords` (type: `string`):

Free-text search query, matching sibling actor hh-ru-jobs-scraper's field name so results can be merged across both. Sent as Habr Career's own search query. Ignored when Division or Skills below is set, since Habr Career only supports one filter dimension per search (see the README for details).

## `vacancyUrls` (type: `array`):

Specific Habr Career vacancy URLs (or bare numeric vacancy IDs) to scrape instead of running a search. When set, this replaces Keywords/Division/Skills/City entirely. Also the source for a repeat check of a known shortlist across runs.

## `division` (type: `string`):

Browse a single specialization/division, using Habr Career's own taxonomy path (group/alias). Real, confirmed slugs are suggested below, but any valid Habr Career division path can be typed in. Takes priority over Skills and Keywords/City when set, since Habr Career's site only supports one filter dimension per search.

## `skills` (type: `array`):

Filter by a single skill tag, using Habr Career's own canonical skill slug (e.g. spring-boot, docker, kubernetes). Real, confirmed slugs are suggested below. Only the first value is used if more than one is given, since Habr Career's skill browse page does not support combining skills. Ignored when Division is set.

## `city` (type: `string`):

Filter to vacancies located in one Russian city. Resolved to Habr Career's own internal city ID via a small built-in lookup of major tech hubs; an unrecognized city is ignored with a warning rather than guessed. Combines with Keywords, but ignored when Division or Skills is set.

## `proxyConfiguration` (type: `object`):

Datacenter proxy is sufficient for Habr Career: real-Chrome, Googlebot, and empty-User-Agent requests all reached full real content with no CAPTCHA or WAF challenge encountered at the depth tested (QRATOR, present in the server header, is confirmed purely passive).

## `includeSalaryBenchmarks` (type: `boolean`):

Fetches Habr Career's own public Russian IT salary survey once per run and adds it as a separate Salary Benchmarks dataset view: real, large-sample compensation data (min/max/percentiles/sample size) broken down by seniority level (Junior/Middle/Senior/Lead). Cheap regardless of Maximum Vacancies, since it is one extra page fetch per run, not per vacancy, so it defaults on.

## `maxItems` (type: `integer`):

Maximum number of vacancies to return per run. Kept low by default so the unmodified default input finishes quickly; raise it for a larger pull once you know your search returns useful results.

## `onlyNewListings` (type: `boolean`):

New-postings monitor mode: re-runs the same search (or Vacancy URLs list) and outputs only vacancies not seen in a previous run that used the same State Name below, a standing-search new-listing alert. Run this Actor on a schedule with the same State Name to build a watch.

## `stateName` (type: `string`):

Identifies which saved search this run's monitor state belongs to. Use a different value for each independent watch you want to track in parallel (e.g. "backend-moscow" vs "devops-remote").

## `resetState` (type: `boolean`):

Clears the persisted monitor state for State Name before this run, so the next run treats every vacancy as new again. Use this to start a watch over from scratch.

## Actor input object example

```json
{
  "keywords": "python",
  "vacancyUrls": [],
  "division": "",
  "skills": [],
  "city": "",
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "includeSalaryBenchmarks": true,
  "maxItems": 20,
  "onlyNewListings": false,
  "stateName": "default",
  "resetState": false
}
```

# Actor output Schema

## `vacancies` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": "python",
    "vacancyUrls": [],
    "division": "",
    "skills": [],
    "city": "",
    "proxyConfiguration": {
        "useApifyProxy": true
    },
    "includeSalaryBenchmarks": true,
    "maxItems": 20,
    "onlyNewListings": false,
    "stateName": "default",
    "resetState": false
};

// Run the Actor and wait for it to finish
const run = await client.actor("getascraper/habr-career-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": "python",
    "vacancyUrls": [],
    "division": "",
    "skills": [],
    "city": "",
    "proxyConfiguration": { "useApifyProxy": True },
    "includeSalaryBenchmarks": True,
    "maxItems": 20,
    "onlyNewListings": False,
    "stateName": "default",
    "resetState": False,
}

# Run the Actor and wait for it to finish
run = client.actor("getascraper/habr-career-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": "python",
  "vacancyUrls": [],
  "division": "",
  "skills": [],
  "city": "",
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "includeSalaryBenchmarks": true,
  "maxItems": 20,
  "onlyNewListings": false,
  "stateName": "default",
  "resetState": false
}' |
apify call getascraper/habr-career-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,getascraper/habr-career-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KsfohE63YWgk5Szwn/builds/BOG8Dfb1eEFTXoPPC/openapi.json
