# Hiring.Cafe Jobs Scraper (`corvuslab/hiringcafe-scraper`) Actor

Scrape hiring.cafe job listings into structured records — title, company, salary, seniority, workplace type, requirements, benefits and enriched company data — with keyword/location/URL search, incremental monitoring and notifications.

- **URL**: https://apify.com/corvuslab/hiringcafe-scraper.md
- **Developed by:** [Corvuslab](https://apify.com/corvuslab) (community)
- **Categories:** Lead generation, Jobs, Automation
- **Stats:** 1 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.15 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Hiring.Cafe Jobs Scraper — Structured Job Data with Salary & Company Info

> **Turn any hiring.cafe search into clean, structured job data — salary, seniority, skills, benefits and enriched company profiles — in seconds, at scale.**

Hiring.Cafe aggregates live job postings from thousands of company career pages and ATS platforms — a single search can surface **250,000+ live listings**. This scraper turns any hiring.cafe search into a rich, typed dataset: job **title, company, location, salary band, seniority, commitment, required skills and languages, benefits, and full descriptions**, plus an **enriched company profile** (HQ country, size, year founded, industries, stock symbol). Search by keyword and location, apply the site's own filters, paste a hiring.cafe search URL, or point it at a single company's page — then export to JSON, CSV, Excel or the API, no code required.

**You pay only for the results you keep** — no proxy required, and incremental mode bills only what changed (see below).

**Why this scraper**

- ⚡ **Fast & low-cost** — a lean, direct data path with no heavy browser, so large runs stay cheap.
- 🧾 **Rich, typed records** — 90+ structured fields per job, not raw HTML.
- ♻️ **Cheap to monitor** — incremental mode re-scrapes only what changed (see below).
- 🔔 **Notifications built in** — Telegram, Slack, Discord or any webhook.
- 🤖 **AI- & API-ready** — compact output, MCP-friendly, one-click integrations.

***

### ✨ Key features

- 🔎 **Search, URL or company scraping** — run a keyword + location + filter search, paste hiring.cafe search URLs (filters and all), or drop a company's hiring.cafe page to pull every open role at that employer.
- 🧾 **Deeply structured records** — normalized salary (min/max/currency/period), seniority, workplace type, job category, required skills, languages, degree requirements and benefit flags.
- 🏢 **Enriched company data** — website, sector, HQ country, employee count, year founded, industries, organization type and stock symbol.
- 🎚️ **Rich filters** — workplace type, commitment, seniority, company, "posted within N days" and sort order, all applied on the site.
- 🧠 **Full descriptions & contacts** — fetch each job's complete description as text, HTML **and** Markdown, with emails/phones auto-extracted.
- ♻️ **Incremental monitoring** — schedule it and get only what changed (NEW / UPDATED / EXPIRED); unchanged jobs are skipped *before* their description is even fetched.
- 🔔 **Notifications** — Telegram, Slack, Discord or any webhook (n8n / Make / Zapier).
- 🤖 **AI-ready** — compact and drop-empty output modes keep payloads small for LLMs and MCP.

***

### 📤 Example output

```json
{
  "id": "dayforce___stretto___strettocareers___152",
  "title": "Software Engineer",
  "company": "Stretto",
  "companyWebsite": "stretto.com",
  "location": "United States or Irvine",
  "workplaceType": "Remote",
  "seniorityLevel": "Mid Level",
  "commitment": ["Full Time"],
  "jobCategory": "Software Development",
  "salaryMin": 110000,
  "salaryMax": 130000,
  "salaryCurrency": "USD",
  "salaryPeriod": "Yearly",
  "salaryRange": "USD 110,000 - USD 130,000 (Yearly)",
  "isCompensationTransparent": true,
  "technicalTools": ["Java", "Spring", "Node.js", "ReactJS", "AWS", "SQL", "Python"],
  "languageRequirements": ["English"],
  "requirementsSummary": "5+ years software development; Bachelor's degree in CS; strong backend and frontend skills; AWS, SQL and CI/CD.",
  "visaSponsorship": false,
  "companyHqCountry": "US",
  "companyYearFounded": 1986,
  "companyIndustries": ["Legal Tech", "Financial Services"],
  "postedDate": "2026-05-14T07:00:00.000Z",
  "applyUrl": "https://jobs.dayforcehcm.com/en-US/stretto/strettocareers/jobs/152",
  "source": "hiring.cafe",
  "description": "Full job description as text, HTML and Markdown…"
}
```

Descriptions are returned as plain **text**, **HTML** and **Markdown** — pick one or all three.

***

### 📚 What data can you extract?

- **Core** — id, title, company, location, workplace type, seniority, commitment, job category, apply URL and source ATS.
- **Compensation** — salary min/max, currency, period, formatted range and a transparency flag.
- **Requirements** — skills/tools, languages, degree requirements by level, licenses/certifications, security clearance, years of experience.
- **Benefits & attributes** — remote/visa/relocation/retirement/PTO/parental-leave flags, shift & travel needs, physical/cognitive demand.
- **Company** — website, sector, HQ country, employee count, year founded, industries, organization type, stock exchange & symbol.
- **Details** (with `includeDetails`) — full description (text / HTML / Markdown) and auto-extracted contact emails & phones.
- **Monitoring** — `changeType`, repost detection, `scrapedAt`, content hash.

Every field is present in standard mode (missing values are `null`); **compact mode** returns the core fields only, for lean AI/MCP payloads.

***

### ⚙️ Input

Configure it in the visual editor — no code needed — or pass JSON via the API.

| Field | What it does |
|---|---|
| `query` | Keyword search, e.g. "registered nurse" (comma-separate for multiple searches). |
| `location` | City, region or country to search near, e.g. "London", "New York". |
| `startUrls` | Paste hiring.cafe search URLs (every filter preserved) or a company's page to pull all its open roles. |
| `workplaceTypes` | Remote, Hybrid, Onsite, Field. |
| `commitmentTypes` | Full Time, Part Time, Contract, Internship, Temporary, Seasonal, Volunteer. |
| `seniorityLevel` | No Prior Experience Required, Entry, Mid, Senior. |
| `dateFetchedPastNDays` | Only jobs added in the last N days. |
| `includeDetails` | Fetch each job's full description and contacts. |
| `incrementalMode` | Emit only what changed since the last run. |
| `maxResults` | Cap the number of jobs (0 = unlimited). |

…and **31 inputs** in total — the table shows the essentials; the rest cover sorting, company & advanced search-state filters, output/AI modes, notification channels and tuning, all in the visual editor.

***

### 📥 Example inputs

Copy any of these into the JSON editor, or set the same fields in the visual editor.

**1. Simple keyword + location search** — grab the 200 newest nursing roles near New York.

```json
{ "query": "registered nurse", "location": "New York", "maxResults": 200 }
```

**2. Filtered search with full descriptions** — senior, remote Python roles, each with its complete description and contacts.

```json
{ "query": "python developer", "workplaceTypes": ["Remote"], "seniorityLevel": ["Senior Level"], "includeDetails": true }
```

**3. Every open job at one company** — paste a company's hiring.cafe page to pull all its roles, enriched.

```json
{ "startUrls": [{ "url": "https://hiring.cafe/org/stripe" }], "includeDetails": true }
```

**4. Paste a search URL** — reuse a hiring.cafe search you built in the browser, filters and all.

```json
{ "startUrls": [{ "url": "https://hiringcafe.com/?searchState=%7B%22searchQuery%22%3A%22data+engineer%22%7D" }] }
```

**5. Scheduled monitoring with alerts** — watch a search daily and get only new/changed jobs pushed to Telegram.

```json
{ "query": "product manager", "location": "London", "incrementalMode": true, "telegramToken": "…", "telegramChatId": "@mychannel" }
```

**6. Lean output for AI/MCP** — core fields only, empties dropped, for small LLM payloads.

```json
{ "query": "solutions architect", "compact": true, "excludeEmptyFields": true }
```

***

### 💡 Use cases

- **Recruiting & sourcing** — build a live, structured feed of open roles by title, location and seniority.
- **Labor-market & salary research** — analyze compensation, skills demand and hiring trends across companies and industries.
- **Company & competitor monitoring** — paste a company's hiring.cafe page to pull every open role, enriched with size, sector and HQ, and watch how its hiring changes over time.
- **Lead generation** — surface companies that are actively hiring, enriched with website, size and sector.
- **Job-board & aggregator building** — power your own board with clean, de-duplicated postings and full descriptions.
- **Vacancy monitoring** — schedule it with incremental mode + notifications for a live change feed of new openings.
- **AI agents & pipelines** — compact output plugs straight into LLM/MCP workflows.

***

### ♻️ Incremental monitoring — pay for changes, not repeats

Schedule the actor and turn on **incremental mode**: each run compares against the last and emits only **NEW / UPDATED / EXPIRED** jobs — unchanged postings are skipped *before* their description is fetched, so a daily watch costs a fraction of a full re-scrape.

| Daily churn | of 1,000 tracked | billable records | you save |
|---|---|---|---|
| 5 % | 1,000 | 50 | **95 %** |
| 15 % | 1,000 | 150 | **85 %** |
| 30 % | 1,000 | 300 | **70 %** |

The first run seeds the baseline and bills in full; every run after that bills only the delta. Reposts of the same role are flagged so you can suppress duplicates.

***

### 🚀 How to run it

1. Open the actor and enter a **search keyword** (e.g. `registered nurse`) and a **location** — or paste a hiring.cafe URL.
2. Pick any **filters** (workplace type, commitment, seniority, posted-within days).
3. Set **Max results** and choose whether to **fetch full descriptions**.
4. (Optional) Turn on **incremental mode** and a **notification** channel, then **Schedule** it.
5. Click **Start**, then download the data as **JSON, CSV or Excel**, or pull it from the **API**.

New to Apify? Create a free account — it comes with monthly credit, no credit card required.

***

### 🔌 Integrations & export

Export to **JSON, CSV, Excel** or an HTML table, or pull from the **REST API** and the **JavaScript / Python** clients. Runs on a **schedule**, connects to **Google Sheets, Slack, Make, Zapier and n8n**, and works as an **MCP tool** for AI agents — compact mode keeps token usage small.

***

### ❓ FAQ

**Do I need a proxy or login?** No — it runs out of the box with no login and no proxy. Apify Proxy is available under Advanced for very high-volume runs.

**Can I scrape all jobs at one specific company?** Yes — paste that company's hiring.cafe page URL (e.g. `hiring.cafe/org/stripe`) into `startUrls` and it pulls every open role, with enriched company data.

**Can I get only new jobs on a schedule?** Yes — turn on incremental mode and schedule it; each run emits only what changed and can notify your channel.

**Can I get remote-only jobs?** Yes — set `workplaceTypes` to `Remote` (combine it with a keyword, location or seniority filter).

**Does it work with n8n, Make or Zapier?** Yes — it runs on a schedule and connects to n8n, Make, Zapier, Google Sheets and Slack, or any webhook.

**How much does it cost?** You pay only for the results you keep — and because incremental mode bills only changed jobs, scheduled monitoring costs a fraction of a full re-scrape.

**Does it get salaries and full descriptions?** Yes — normalized salary bands (where the employer discloses them) and full descriptions as text, HTML and Markdown.

**What formats can I export?** JSON, CSV, Excel, HTML table, or via the API.

**Is it good for AI agents?** Yes — enable compact mode; the output is MCP-friendly.

**How many jobs can I get?** As many as the search returns — set `maxResults` (0 = unlimited).

**Is scraping this legal?** The actor collects only **publicly available** data. You are responsible for how you use it, including any personal data and GDPR-style obligations.

***

### ⚖️ Disclaimer

This actor accesses only publicly available data on hiring.cafe. You are responsible for how you use the extracted data — in particular any personal information — and for complying with the site's terms and applicable law (including the GDPR where it applies). Not affiliated with, endorsed by, or sponsored by hiring.cafe.

***

**Keywords:** hiring.cafe scraper · hiring cafe scraper · hiringcafe api · hiring.cafe job scraper · job scraper · job listings api · job data extraction · salary data scraper · remote jobs scraper · ATS jobs scraper · scrape jobs by company · company hiring data · company jobs scraper · recruitment data · labor market research · job board scraper · export jobs to CSV Excel JSON · no-code job scraper · MCP tool for AI agents

# Actor input Schema

## `query` (type: `string`):

Job title or keyword, e.g. "registered nurse" or "python developer". Separate multiple searches with commas — each runs as its own search and results are merged and de-duplicated.

## `location` (type: `string`):

City, region or country to search near, e.g. "London", "New York", "Germany". Resolved to the site's nearest match. Leave empty for a location-agnostic search (great with Remote).

## `startUrls` (type: `array`):

Paste hiring.cafe search URLs (the ones with all your filters applied, e.g. https://hiringcafe.com/?searchState=...) or company pages. Every filter in the URL is preserved.

## `maxResults` (type: `integer`):

Maximum number of jobs to return. Set 0 for unlimited (bounded by how many the search has).

## `ignoreUrlFailures` (type: `boolean`):

Skip start URLs that cannot be interpreted instead of failing the whole run.

## `workplaceTypes` (type: `array`):

Remote, hybrid, onsite or field-based roles.

## `commitmentTypes` (type: `array`):

Employment type.

## `seniorityLevel` (type: `array`):

Experience level required.

## `companyKeywords` (type: `string`):

Restrict to companies matching these keywords (comma-separated), e.g. "Siemens", "Google".

## `dateFetchedPastNDays` (type: `integer`):

Only jobs added in the last N days. Leave empty (or 0) for any time.

## `sortBy` (type: `string`):

Result ordering.

## `searchStateOverrides` (type: `object`):

Power users: extra raw search-state keys merged on top of the filters above (e.g. salary bounds, industries, benefits). Leave empty unless you know the field names.

## `includeDetails` (type: `boolean`):

Fetch each job's full description and extract contacts. Turn off for the fastest, cheapest runs (the structured fields still come through).

## `descriptionFormat` (type: `string`):

Which representation(s) of the job description to include.

## `descriptionMaxLength` (type: `integer`):

Truncate each description to this many characters. Leave empty for the full text.

## `compact` (type: `boolean`):

Emit only the core fields (title, company, location, salary, apply URL, …). Ideal for AI agents and MCP clients.

## `excludeEmptyFields` (type: `boolean`):

Remove null, empty-string and empty-array fields from each record.

## `incrementalMode` (type: `boolean`):

Track state between runs and tag every job with a changeType (NEW / UPDATED / UNCHANGED / EXPIRED), with repost detection.

## `stateKey` (type: `string`):

Stable name for the tracked search. Leave empty to derive one automatically from your search settings.

## `emitUnchanged` (type: `boolean`):

Also emit jobs that have not changed since the previous run.

## `emitExpired` (type: `boolean`):

Emit jobs that were present last run but are gone now.

## `telegramToken` (type: `string`):

Bot token from @BotFather.

## `telegramChatId` (type: `string`):

Chat or channel ID, e.g. "-100123456789" or "@yourchannel".

## `slackWebhookUrl` (type: `string`):

Slack incoming-webhook URL.

## `discordWebhookUrl` (type: `string`):

Discord incoming-webhook URL.

## `webhookUrl` (type: `string`):

Any HTTPS endpoint. Receives a JSON POST with the matched jobs — works with n8n, Make and Zapier.

## `webhookHeaders` (type: `object`):

Extra headers for the webhook request, e.g. {"Authorization": "Bearer xyz"}.

## `notificationLimit` (type: `integer`):

How many jobs to include in each notification message.

## `proxyConfiguration` (type: `object`):

Optional. The site works over a direct connection, so no proxy is used by default. Enable Apify Proxy only if you run very high volume and want to spread requests across IPs.

## `detailConcurrency` (type: `integer`):

How many full descriptions to fetch in parallel.

## `maxRequestRetries` (type: `integer`):

How many times to retry a failed request before giving up on it.

## Actor input object example

```json
{
  "query": "registered nurse",
  "maxResults": 25,
  "ignoreUrlFailures": true,
  "sortBy": "relevance",
  "includeDetails": true,
  "descriptionFormat": "all",
  "compact": false,
  "excludeEmptyFields": false,
  "incrementalMode": false,
  "emitUnchanged": false,
  "emitExpired": false,
  "notificationLimit": 5,
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "detailConcurrency": 8,
  "maxRequestRetries": 3
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `allItems` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "python developer",
    "location": ""
};

// Run the Actor and wait for it to finish
const run = await client.actor("corvuslab/hiringcafe-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "query": "python developer",
    "location": "",
}

# Run the Actor and wait for it to finish
run = client.actor("corvuslab/hiringcafe-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "python developer",
  "location": ""
}' |
apify call corvuslab/hiringcafe-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,corvuslab/hiringcafe-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XVCGw7l2bJuxcxxHG/builds/8YGDerjuOmduhwgAs/openapi.json
