# HN Who is Hiring — AI Structured Jobs (`arthur_taken/hn-who-is-hiring-ai`) Actor

Hacker News monthly 'Who is hiring?' posts parsed by AI into clean job rows: salary range, remote regions, visa, tech stack, apply links.

- **URL**: https://apify.com/arthur\_taken/hn-who-is-hiring-ai.md
- **Developed by:** [DH](https://apify.com/arthur_taken) (community)
- **Categories:** Jobs, AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.80 / 1,000 job rows

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## HN Who is Hiring — AI Structured Jobs

Every month, hundreds of companies post jobs in Hacker News' **"Ask HN: Who is hiring?"** thread. The posts are free-form text, so salary, remote rules, visa policy and tech stack are hard to filter.

This Actor uses AI (Claude) to turn each post into **clean, filterable job rows** — one row per role, even when a single post lists ten roles.

### What you get

| Field | Example |
|---|---|
| `company_name`, `company_summary` | Discord · "Communities and friend groups spend time together…" |
| `title`, `seniority`, `category` | Senior Software Engineer, Application Security · `senior` · `security` |
| `work_mode`, `locations`, `remote_regions`, `timezone_note` | `onsite` · San Francisco, CA |
| `salary_min`, `salary_max`, `salary_currency`, `salary_period`, `salary_is_total_comp`, `has_equity` | 280000 · 330000 · USD · year · true |
| `visa_sponsorship`, `citizenship_required`, `security_clearance` | `unspecified` · null · null |
| `employment_type` | `full_time` / `part_time` / `contract` / `internship` |
| `tech_stack` | Python, TypeScript, React, Rust, Google Cloud Platform |
| `apply_url`, `careers_url`, `company_url`, `contact_emails` | links copied from the post |
| `hn_url`, `posted_at`, `month` | link back to the original comment |
| `job_id` | stable ID per role (`<comment id>-<index>`) for deduplication |
| `first_seen_month`, `is_repost` | whether the same company posted the same title in an earlier month |

Rules the AI follows:

- **No guessing.** If the post doesn't say it, the field is `null` / `unspecified`. A plain "REMOTE" is not turned into "Worldwide".
- **Links are verified.** Every URL is checked against the original text; links that don't appear in the post are removed.
- **Multi-role posts are split** into one row per role, with shared salary/location copied only where the post applies it.

### Why AI instead of regex

Regex parsers depend on the `Company | Role | Location` first line and return empty fields when a poster writes a paragraph instead. They also miss salary formats like `100-200k CHF` or `$280K–$330K TC`. AI reads the whole post the way a person would.

### Input

| Option | Description |
|---|---|
| `months` | `latest`, `all`, or `YYYY-MM` (e.g. `2026-09`). Several months allowed. |
| `workModes` | `remote`, `hybrid`, `onsite`, `remote_or_onsite`. For remote jobs use both `remote` and `remote_or_onsite`. |
| `techStack` | any of, exact name (case-insensitive), e.g. `Python`, `PostgreSQL`, `Node.js`. Add aliases when unsure (`Rails`, `Ruby on Rails`). |
| `regions` | any of, matched at word start in remote regions and office locations, e.g. `EU`, `Berlin`, `Worldwide` |
| `keywords` | any of, matched at word start in title, company, summary, tech stack (`engineer` matches `Engineering`) |
| `requireVisaSponsorship` | only posts that explicitly offer sponsorship |
| `excludeCitizenshipRequired` | drop roles requiring citizenship or clearance |
| `seniority`, `categories`, `employmentTypes` | any of the listed values, e.g. `senior`, `data`, `full_time` |
| `minAnnualSalary`, `salaryCurrencies` | e.g. `150000` and `USD`. Hourly/monthly pay is annualized. Roles without a salary are dropped. |
| `postedAfter` | `2026-09-15` or relative like `7 days` |
| `excludeReposts` | hide roles the company already posted in an earlier month |
| `onlyNew` | for scheduled runs: return only rows you haven't received yet with the same filters |
| `maxItems` | cap the number of rows (and your cost) |

Example — remote Python roles open to the EU:

```json
{
  "months": ["latest"],
  "workModes": ["remote"],
  "techStack": ["Python"],
  "regions": ["EU", "Europe", "Worldwide"]
}
```

### Using it from an AI agent (MCP)

The Actor works well as a tool for AI agents via the [Apify MCP server](https://mcp.apify.com):

- A minimal call: `{"months": ["latest"], "workModes": ["remote", "remote_or_onsite"], "techStack": ["Python"], "maxItems": 50}`
- Every row has the same fields; missing information is `null` or `"unspecified"`, never guessed.
- If nothing matches, the run's status message says so and lists the months that have data.
- Set `maxItems` (or a spending limit) to control cost — the run stops cleanly at your limit.

### Coverage and freshness

- **Months covered:** every monthly thread from **October 2025** onward (12+ months), and each new month is added automatically.
- **Freshness:** new threads are processed within a day of posting; late comments are added daily. Right after a new thread appears, `latest` keeps returning the previous month until the new one is ready.
- **Tech names are unified:** `Postgres` → `PostgreSQL`, `NextJS` → `Next.js`, `Rails` → `Ruby on Rails`, so `techStack` filters don't miss aliases.

### Daily job alerts (scheduled runs)

Schedule this Actor daily with `onlyNew: true` and your filters to get only fresh roles each day, e.g.:

```json
{
  "months": ["latest"],
  "workModes": ["remote", "remote_or_onsite"],
  "techStack": ["Python", "PostgreSQL"],
  "excludeReposts": true,
  "onlyNew": true
}
```

You pay only for new rows. Job IDs you already received are remembered in a key-value store named `hn-who-is-hiring-state` in your own account.

### Pricing

**$1.80 per 1,000 job rows.** You pay only for rows that match your filters. Use `maxItems` to cap spend.

Runs are fast (a few seconds) because posts are parsed once in advance — your run only filters.

### Output sample

```json
{
  "month": "2026-09",
  "company_name": "Discord",
  "title": "Senior Software Engineer, Application Security",
  "seniority": "senior",
  "category": "security",
  "employment_type": "full_time",
  "work_mode": "onsite",
  "locations": ["San Francisco, CA"],
  "remote_regions": [],
  "salary_min": 280000,
  "salary_max": 330000,
  "salary_currency": "USD",
  "salary_period": "year",
  "salary_is_total_comp": true,
  "visa_sponsorship": "unspecified",
  "tech_stack": ["Python", "TypeScript", "React", "Rust", "JavaScript", "Google Cloud Platform"],
  "apply_url": "https://job-boards.greenhouse.io/discord/jobs/8656854002",
  "hn_url": "https://news.ycombinator.com/item?id=49525748"
}
```

### Data source and responsible use

- Data comes from the official Hacker News API (public posts in the monthly "Who is hiring?" thread).
- `contact_emails` contains only addresses written out in plain text by the poster. Obfuscated addresses (e.g. `name [at] domain [dot] com`) are **not** decoded.
- Use contact details only to apply for the posted roles. Don't use them for bulk marketing. You are responsible for complying with applicable laws such as GDPR.
- AI extraction is very accurate but not perfect. Check the original post (`hn_url`) before acting on salary or visa details.

### FAQ

**Is "Who wants to be hired?" included?** Not yet — this Actor covers the hiring thread only.

**How are salaries normalized?** `60-100k` → `60000`–`100000`. Currency is ISO 4217. `TC` sets `salary_is_total_comp: true`. Hourly rates use `salary_period: "hour"`.

**Found a wrong row?** Open an issue with the `hn_url` and we'll fix the parser.

### Related

- [HN Who wants to be hired — AI Structured Candidates](https://apify.com/arthur_taken/hn-who-wants-to-be-hired): the candidates' side — developers looking for work, with seniority, years, location, remote/relocation, visa, tech stack, email and résumé.

# Actor input Schema

## `months` (type: `array`):

Which monthly threads to search: 'latest' (newest month with data), 'all' (every month covered, about the last 12), or YYYY-MM such as '2026-09'. More months = more rows and more cost. If a month has no data, the run log and status message list the months that do.

## `workModes` (type: `array`):

Keep only these work modes. For remote jobs, include both 'remote' and 'remote\_or\_onsite' (roles where the candidate can choose). Empty = all modes, including 'unspecified'.

## `techStack` (type: `array`):

Keep roles whose tech\_stack contains at least one of these names. Exact name match, case-insensitive, so use canonical names: 'PostgreSQL' (not 'Postgres'), 'Node.js', 'Go', 'Kubernetes'. Add aliases when unsure, e.g. \['Rails', 'Ruby on Rails'] or \['GCP', 'Google Cloud Platform']. For looser text search use keywords.

## `regions` (type: `array`):

Keep roles whose remote regions or office locations contain one of these terms (case-insensitive, matched at word start). Common remote regions: US, Worldwide, EU, Europe, Canada, UK, LATAM. Office locations are free text such as 'Berlin, Germany' or 'Austin, TX', so 'US' does not match 'Austin, TX' — add city or state names if you need onsite roles.

## `keywords` (type: `array`):

Free-text filter over job title, company name, company summary and tech stack (case-insensitive, matched at word start: 'engineer' matches 'Engineering', 'ML' does not match 'HTML'). Use it for roles or domains, e.g. 'data engineer', 'infrastructure', 'fintech'.

## `requireVisaSponsorship` (type: `boolean`):

Keep only roles whose post explicitly offers visa sponsorship (visa\_sponsorship = 'yes'). Most posts don't mention visas, so this removes about 98% of rows.

## `excludeCitizenshipRequired` (type: `boolean`):

Drop roles that state a citizenship/residency requirement or require a security clearance. Useful for candidates outside the hiring country.

## `seniority` (type: `array`):

Keep only these seniority levels. Most posts don't state seniority ('unspecified' is about 40% of rows), so include 'unspecified' if you don't want to miss roles.

## `categories` (type: `array`):

Keep only these job families (inferred from the title and duties), e.g. 'data', 'ml\_ai', 'devops\_infra'.

## `employmentTypes` (type: `array`):

Keep only these contract types. Most roles are 'full\_time'.

## `minAnnualSalary` (type: `integer`):

Keep roles whose stated salary (upper bound, or lower bound if no upper bound) is at least this much per year. Hourly and monthly pay are converted (x2080, x12). Roles without a stated salary (about 70%) are dropped. Combine with salaryCurrencies to avoid mixing currencies.

## `salaryCurrencies` (type: `array`):

Keep roles paid in these ISO 4217 currencies, e.g. USD, EUR, GBP. Roles without a stated salary are dropped when this is set.

## `postedAfter` (type: `string`):

Keep posts published on or after this date: '2026-09-15' or relative like '7 days'. Applies within the selected months, so use months \['all'] for ranges that cross a month.

## `excludeReposts` (type: `boolean`):

Drop roles that the same company already posted with the same title in an earlier month (is\_repost = true). Useful to see only genuinely new openings.

## `onlyNew` (type: `boolean`):

For scheduled runs: return only job rows you haven't received from this Actor with the same filters. Seen job IDs are kept in a key-value store named 'hn-who-is-hiring-state' in your account. With months 'latest', the last 2 months are checked so late posts are not missed.

## `maxItems` (type: `integer`):

Maximum number of job rows to return. Each row is billed, so this caps your cost. Empty = all matching rows (about 500-650 per month when no other filter is set).

## Actor input object example

```json
{
  "months": [
    "2026-09",
    "2026-08"
  ],
  "workModes": [
    "remote",
    "remote_or_onsite"
  ],
  "techStack": [
    "Python",
    "PostgreSQL"
  ],
  "regions": [
    "EU",
    "Europe",
    "Worldwide"
  ],
  "keywords": [
    "data engineer",
    "platform"
  ],
  "requireVisaSponsorship": false,
  "excludeCitizenshipRequired": false,
  "seniority": [
    "senior",
    "staff",
    "unspecified"
  ],
  "categories": [
    "data",
    "ml_ai"
  ],
  "employmentTypes": [
    "full_time"
  ],
  "minAnnualSalary": 150000,
  "salaryCurrencies": [
    "USD"
  ],
  "postedAfter": "7 days",
  "excludeReposts": false,
  "onlyNew": false,
  "maxItems": 100
}
```

# Actor output Schema

## `jobs` (type: `string`):

All matching job rows (company, title, work mode, salary, visa, tech stack, apply link).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "months": [
        "latest"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("arthur_taken/hn-who-is-hiring-ai").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "months": ["latest"] }

# Run the Actor and wait for it to finish
run = client.actor("arthur_taken/hn-who-is-hiring-ai").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "months": [
    "latest"
  ]
}' |
apify call arthur_taken/hn-who-is-hiring-ai --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,arthur_taken/hn-who-is-hiring-ai"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bzpdpUaayDLCEbFX3/builds/iAhP8LXwD2ghiUeAJ/openapi.json
