# Hacker News Who Is Hiring: Jobs & Hiring History (`datagrit/hacker-news-hiring-history`) Actor

Turn Hacker News Who is hiring threads into job rows with salary, stack, work mode and visa, plus how many months each company has been hiring.

- **URL**: https://apify.com/datagrit/hacker-news-hiring-history.md
- **Developed by:** [datagrit](https://apify.com/datagrit) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Hacker News Who Is Hiring: Jobs & Hiring History do?

Hacker News Who Is Hiring: Jobs & Hiring History turns the monthly "Ask HN: Who is hiring?" thread into a clean table of job posts. Every top-level comment becomes one row with the company, roles, location, work mode, employment type, stated pay, visa sponsorship, technologies, apply link and the full post text. What other scrapers of this thread do not give you is company history: the Actor compares every company with the earlier monthly threads and tells you how many months it has been hiring, when it first and last posted before, and whether it is posting for the first time. It reads the public Hacker News search API, needs no login or API key, and exports to JSON, CSV or Excel.

### Who is it for?

- **Job seekers and recruiters** who want the current thread filtered by technology, location, remote work, salary or visa sponsorship instead of scrolling 400 comments.
- **Sales and lead generation teams** selling to startups: companies posting for the first time are young teams that just got funded or started hiring, and companies that post every month are in steady growth.
- **Market analysts and investors** tracking which technologies and salary levels appear in startup hiring over many months.
- **Developers and data teams** who need a scheduled feed that delivers only the posts they have not seen yet.

### How to use it

1. Choose which thread to read: the latest one, the last few months, or one specific month such as `2026-03`.
2. Set how many earlier months each company is compared against. Twelve months is the default. Zero switches history off and makes the run faster.
3. Add filters if you want a narrower list (below), set the maximum number of posts and run the Actor.

#### Filters

- **Post text contains**, **Role contains**, **Location contains**: keep posts matching any of your words. Role and location are searched in the post header, where companies list them.
- **Technologies**: keep posts naming at least one of the listed technologies, such as python, rust or kubernetes.
- **Work mode** and **Employment type**: remote, hybrid or onsite; full-time, part-time, contract or internship.
- **Only posts with a salary** and **Minimum yearly salary** with a currency: pay is converted to a yearly figure, so hourly and monthly rates compare with annual ones.
- **Only with visa sponsorship**: posts that offer visa or work-permit sponsorship. Posts that refuse it are never returned by this filter.
- **Only first-time companies** and **Minimum months hiring**: use the history to find new teams or companies that hire all the time.
- **Only posts not delivered before**: skip every post an earlier run with the same filters already returned, so a scheduled run delivers each new post once.

#### Company history

For each post the Actor normalises the company name (lower case, letters and digits only, legal suffixes removed) and, when the post links the company's own website, also keeps the domain (`acme.com` for Acme, `factory.ai` for Factory AI). It then looks for the name or the domain in the earlier threads, so a company that links its site in one month and not in the next is still one company. A link to a news article, a LinkedIn short link or a file host inside the post text is never taken for the company's domain. The `companyKey` column is the normalised name alone (`acme`, `factoryai`), so it does not change with the link; two spellings of one company (`Factory` and `Factory AI`) give two keys, and the history columns still treat them as one when the website domain is the same. You get `monthsHiring` (this month included), `firstSeenMonth`, `previousSeenMonth`, `consecutiveMonths`, `isNewCompany` and `postsThisMonth`. Posts where no company could be read from the header, or only a placeholder such as `Stealth`, `Confidential` or `Startup`, have no key and carry null history, `monthsChecked` included, rather than a guess. A header that starts with a place (`Boston, MA - Full Time`, `NYC | Engineer`; a name with a place after a comma, such as `adapptiv Labs, Switzerland` or `LiquidFi (Miami, FL)`, stays a company), a role (`Software Engineer | Acme`) or an employment phrase gives no company either, and those rows are delivered free of charge, as are rows with a placeholder name. Older threads (before about 2016) were written free-form, so many of their posts have no readable company.

History is only computed from threads whose posts can be read reliably: every thread of the window, and the thread you ask for, must have a standard `Company | role | location` header in at least 75% of its posts. Before 2016 that is not the case (see "Good to know"), so a run that asks for history on such a month, or on a month whose window reaches one, stops with a message instead of returning rows with empty history. Set **Months of history** to 0 to read those months as plain posts: you then get the posts without history fields, and rows without a company name are free.

### Example output

```json
{
  "postId": "49156683",
  "sourceUrl": "https://news.ycombinator.com/item?id=49156683",
  "threadMonth": "2026-09",
  "company": "Modash.io",
  "companyWebsite": "https://modash.io",
  "roles": ["Senior Product Engineer"],
  "location": "Remote (Europe)",
  "workModes": ["remote"],
  "employmentTypes": ["full-time"],
  "salaryMin": 75000,
  "salaryMax": 110000,
  "salaryCurrency": "EUR",
  "salaryPeriod": "year",
  "visaSponsorship": true,
  "technologies": ["TypeScript", "Python"],
  "monthsChecked": 13,
  "monthsHiring": 4,
  "firstSeenMonth": "2026-05",
  "previousSeenMonth": "2026-08",
  "isNewCompany": false
}
```

The values above show the shape of a row. Fields that a post does not state, such as salary, are null instead of an empty string.

### Pricing

You pay per job post returned. Posts removed by your filters, posts skipped as already delivered, posts without a usable company name, placeholders such as `Stealth` included (the run status counts them) and a run that finds nothing are not charged: a search without matches returns a single status row with `found: false` that is free. Reading earlier threads for the history costs nothing extra. Set a maximum charge per run in the Apify console to cap spending; the Actor stops cleanly when it is reached.

### Good to know

- Posts are free text written by people, so parsing is heuristic. The status message at the end of every run reports how many posts had a readable header, how many stated a salary and how many months of the history window were read. It names the threads the posts came from; when **Maximum posts** is reached before the later threads of a multi-month request, it lists those threads as not read, so a missing month is never mistaken for a month without matches. About three in ten recent posts state pay.
- The share of posts with a standard header depends on the year of the thread, because the `Company | role | location` format only settled around 2016. Measured on every thread in October 2026: 2011 and 2012 between 2% and 27%, 2013 between 10% and 35%, 2014 between 22% and 41%, 2015 from 35% in January to 77% in December, 2016 between 79% and 92%, 2017 to 2026 between 89% and 96%. With history switched off, posts without a company name are returned but free. With history on, a thread below 75% stops the run (see Company history); with the default 12 months of history that means asking for a month before December 2016 needs **Months of history** set to 0.
- If Hacker News changes the thread format so that no header can be read in a recent thread, the run fails instead of returning rows with empty fields.
- Company history is only returned when every month of the window is accounted for. The run status says, for example, `History window 12 months (12 read)`. The Actor lists threads from the official whoishiring accounts and then looks up by title every month that list does not contain, or contains only as a thread with fewer than 20 comments, and keeps the thread with the most comments, so threads posted from another account (July 2011) and stub threads next to the real one (December 2011) are read correctly. A month is reported as `with no Hacker News thread` only when that title search finds nothing; such a month is skipped without breaking `consecutiveMonths`. If a month cannot be confirmed, the run fails instead of marking companies as new by mistake, and running it again fixes it.
- The newest thread is skipped while it has fewer than 10 comments, which happens in the first hours of a month. Ask for a specific month to read it anyway.
- Only public posts are read. No account, cookies or private data are involved.

### FAQ

**Is it legal to scrape Hacker News?** The Actor reads the public Algolia search API that Hacker News itself links to, for public job posts. It collects no data behind a login.

**How often is the thread published?** Once a month, on the first working day. Schedule the Actor weekly with "Only posts not delivered before" switched on to collect the new posts as they appear during the month.

**How far back does the history go?** Up to 36 earlier months, as long as those threads are readable (see Good to know). Each earlier thread adds about two seconds to the run.

**Why is a company missing from the history?** Its header gave no company name, or it posted under a different name. Rows without a company key have null history.

**The run failed with HTTP 429 or a message about the waiting budget.** The Hacker News search API limits requests per IP address and the Actor shares an IP address with other runs on the platform. Failed attempts and pauses between attempts draw from one 75-second budget for the whole run, and a request that gets no answer is cut off after 25 seconds (also when a proxy accepts the connection and stays silent). When the budget is used up the Actor stops with an error instead of running past the five-minute limit of a scheduled test; successful requests do not use it. Run again in a few minutes, or switch on the proxy in the input.

**Something looks wrong in the data.** Open an issue on the Actor page with the input you used and the run link.

### Related Actors

Use this Actor together with other job and company data tools from the same publisher, such as the job board scrapers, to compare Hacker News hiring with the careers pages of the same companies.

# Changelog

This Actor's version history is a separate document: https://apify.com/datagrit/hacker-news-hiring-history/changelog.md

# Actor input Schema

## `monthsBack` (type: `integer`):

How many of the most recent monthly threads to read, newest first. 1 is the current month. Ignored when Month is set. A thread with fewer than 10 comments (the first hours of a month) is skipped.

## `month` (type: `string`):

Read one month instead, written as YYYY-MM, for example 2026-03. The run fails when neither the whoishiring accounts nor a title search on Hacker News finds a Who is hiring thread for that month. Before 2016 most posts have no standard header, so such a month needs Months of history set to 0.

## `historyMonths` (type: `integer`):

How many earlier monthly threads to compare every company against, for months hiring, first seen, previous seen and new company. 0 switches history off and makes the run faster; it is also the setting for months before December 2016, where company history cannot be computed and a run with history fails. More months mean more threads to download: each one adds about a second.

## `keywords` (type: `array`):

Keep posts whose full text contains at least one of these words or phrases (case-insensitive), for example climate or "machine learning".

## `roleContains` (type: `array`):

Keep posts whose header line contains at least one of these words, for example designer or "data engineer". The header is where roles are listed.

## `locationContains` (type: `array`):

Keep posts whose header line contains at least one of these words, for example berlin, europe or "new york".

## `technologies` (type: `array`):

Keep posts that name at least one of these technologies anywhere in the text, for example python, rust or kubernetes. Names are matched against the Technologies column, case-insensitive.

## `workMode` (type: `string`):

Keep posts whose header mentions this mode. Posts that do not state a mode are dropped when you choose one.

## `employmentType` (type: `string`):

Keep posts whose header mentions this type. Posts that do not state one are dropped when you choose one.

## `onlyWithSalary` (type: `boolean`):

Keep only posts that state pay in the header or in a Salary line.

## `minSalary` (type: `integer`):

Keep posts whose top stated pay, converted to a year, is at least this amount. Needs Salary currency, because pay is posted in different currencies. 0 means no minimum.

## `salaryCurrency` (type: `string`):

Currency that Minimum yearly salary is measured in. Posts in other currencies are dropped when a minimum is set.

## `visaSponsorshipOnly` (type: `boolean`):

Keep only posts that offer visa or work-permit sponsorship.

## `onlyNewCompanies` (type: `boolean`):

Keep only companies that did not post in any of the earlier threads covered by Months of history. Needs history.

## `minMonthsHiring` (type: `integer`):

Keep companies that posted in at least this many of the checked threads, this one included. Use it to find companies that hire all the time. Needs history. 0 means no minimum.

## `onlyNewSinceLastRun` (type: `boolean`):

Return only posts that no earlier run with the same filters has delivered. Schedule the Actor during the month to get each new post once. Runs with different filters keep separate memory.

## `maxItems` (type: `integer`):

Stop after this many posts in total across all threads read. A monthly thread holds 250 to 450 posts.

## `proxyConfiguration` (type: `object`):

Optional proxy. Leave disabled unless the Hacker News search API rate-limits the shared platform IP (HTTP 429); a proxy raises the platform cost of the run.

## Actor input object example

```json
{
  "monthsBack": 1,
  "historyMonths": 12,
  "workMode": "any",
  "employmentType": "any",
  "onlyWithSalary": false,
  "minSalary": 0,
  "salaryCurrency": "any",
  "visaSponsorshipOnly": false,
  "onlyNewCompanies": false,
  "minMonthsHiring": 0,
  "onlyNewSinceLastRun": false,
  "maxItems": 50,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

All job posts as a dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "monthsBack": 1,
    "historyMonths": 12,
    "maxItems": 50,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("datagrit/hacker-news-hiring-history").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "monthsBack": 1,
    "historyMonths": 12,
    "maxItems": 50,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("datagrit/hacker-news-hiring-history").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "monthsBack": 1,
  "historyMonths": 12,
  "maxItems": 50,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call datagrit/hacker-news-hiring-history --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datagrit/hacker-news-hiring-history"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HIblLaf31mAl5RK0G/builds/5GHN31ZsGBL3orXxq/openapi.json
