# Hacker News Scraper – Stories, Comments, Search & Who's Hiring (`gazidev/hacker-news-data`) Actor

Hacker News data from the official API and Algolia HN Search: top/new/best/Ask/Show/job stories, keyword and brand monitoring with only-new alerts, full comment trees flattened with depth, and the monthly Who is hiring thread parsed into job posts.

- **URL**: https://apify.com/gazidev/hacker-news-data.md
- **Developed by:** [Cemal Atakli](https://apify.com/gazidev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.20 / 1,000 hn items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Hacker News Scraper – Stories, Comments, Search & Who's Hiring

Get **Hacker News data as clean JSON**: front page and new/best/Ask HN/Show HN/job stories, **keyword and brand monitoring** across stories and comments, **full comment trees** flattened with depth and parent IDs, and the monthly **"Ask HN: Who is hiring?" thread parsed into job posts** (company, roles, location, remote, salary, links).

It uses only the **official Hacker News API** (Firebase) and the public **Algolia HN Search API**. No browser, no proxies, no login, so it is fast and costs **$0.20 per 1,000 items**.

### What it does

- **Story lists.** `top` (front page ranking), `new`, `best`, `ask`, `show`, `job`, with title, URL, domain, points, comment count, author, time, text and `rank`.
- **Search & brand monitoring.** Search stories, comments or both through Algolia HN Search. Filter by date range (`2026-09-01` or `7 days`), min points, min comments, author, Show HN / Ask HN / front page. Sort by date (no 1,000-result cap: the Actor pages by time) or relevance. **Strict matching** removes Algolia's typo matches (`rust` → `trust`).
- **Only-new mode for alerts.** Schedule it hourly: each run returns only stories, comments or job posts that were not returned before. Works for searches, lists, comment threads and the hiring thread.
- **Full comment trees.** Give item IDs or URLs and get the item plus every comment as one row each, with `depth`, `parentId`, `rankAmongSiblings`, `replyCount`, plain text (HTML converted, links kept in `links`). Limit depth or count. Top-level comments follow HN's own ranking.
- **"Who is hiring?" parser.** Finds the thread for any month (`latest`, `2026-09`, `March 2019`), splits it into job posts and extracts best-effort fields from the `Company | Role | Location | Remote | Salary | URL` header: `company`, `ycBatch`, `roles`, `locations`, `remote`, `workArrangement`, `employmentType`, `salaryText` / `salaryMin` / `salaryMax` / `salaryCurrency`, `visaSponsorship`, `companyUrl`, `applyUrls`, `emails`. The **full raw text is always kept**. Filter by keywords (`python`, `Berlin`) or remote only.
- Robust: retries with backoff, per-item errors become free `error` rows, honest User-Agent, max 4 parallel requests to Algolia.

### Use cases

- **Brand and competitor monitoring:** get an alert (Slack, email, webhook) whenever your product, domain or competitor is mentioned in a story or comment.
- **Job search and recruiting data:** every Who is hiring post as structured rows; track remote jobs, salaries and stacks month by month.
- **Market and trend research:** what HN says about Rust, Postgres or AI agents over time; top stories by points for a date range.
- **AI / LLM datasets:** clean comment threads with depth and parent IDs for summarisation, sentiment analysis or RAG.
- **Launch tracking:** follow your Show HN thread and get only the new comments each hour.
- **Newsletters and dashboards:** the daily front page or best stories as JSON, CSV or Excel.

### Input examples

Front page (default, one click):

```json
{ "mode": "lists", "lists": ["top"], "maxItemsPerList": 30 }
```

Brand monitoring, scheduled hourly:

```json
{
  "mode": "search",
  "searchQueries": ["Apify", "\"web scraping\"", "crawlee"],
  "searchTags": "storyOrComment",
  "createdAfter": "7 days",
  "onlyNew": true
}
```

A thread with all comments:

```json
{ "mode": "items", "itemIds": ["https://news.ycombinator.com/item?id=8863"], "maxCommentDepth": 0 }
```

Remote Python jobs from this month's Who is hiring:

```json
{ "mode": "whoIsHiring", "hiringMonth": "latest", "hiringKeywords": ["python"], "remoteOnly": true }
```

| Field | Default | Notes |
|---|---|---|
| `mode` | `lists` | `lists`, `search`, `items`, `whoIsHiring` |
| `lists` / `maxItemsPerList` | `["top"]` / 30 | top, new, best, ask, show, job |
| `searchQueries` | – | One per line; quotes for phrases. Rows get `matchedQuery` |
| `searchTags` | `story` | story, comment, storyOrComment, showHn, askHn, poll, frontPage, any |
| `sortBy` | `date` | `relevance` is capped at 1,000 results per query by Algolia |
| `createdAfter` / `createdBefore` | – | `2026-09-01` or `7 days` |
| `minPoints` / `minComments` | 0 | Stories only |
| `author` | – | HN username; works without a query |
| `maxResultsPerQuery` | 100 | |
| `exactMatch` | true | No typo matches; every word must start a word in the text |
| `itemIds` | – | IDs or item URLs (items mode) |
| `includeComments` | false | Also comment trees for list / search stories |
| `maxCommentDepth` / `maxCommentsPerItem` | 0 / 0 | 0 = all |
| `commentSource` | `auto` | `auto` = Algolia tree in one request (falls back to the official API), `firebase` = official API only |
| `hiringMonth` | `latest` | `2026-09`, `September 2026` |
| `hiringKeywords` / `hiringKeywordsMode` / `remoteOnly` | – / any / false | Filters for job posts |
| `onlyNew` / `stateKey` | false / – | Monitoring memory |
| `maxItems` | 0 | Total cap on rows |
| `includeHtml` | false | Adds HN's original HTML as `textHtml` |

### Output examples

One dataset row per item. The Output tab has **Stories & search results**, **Comments**, **Who is hiring – job posts** and **Errors** views. Full examples are in `SAMPLE_OUTPUT.json`.

Story (lists mode):

```json
{
  "type": "story", "id": 49940394, "list": "top", "rank": 1,
  "title": "Newgrounds.com – A community of games, music, and art",
  "url": "https://www.newgrounds.com/", "domain": "newgrounds.com",
  "points": 211, "numComments": 55, "author": "azhenley",
  "createdAt": "2026-10-03T00:55:25Z", "createdAtUnix": 1790988925,
  "parentId": null, "replyCount": 25, "text": null, "links": [],
  "hnUrl": "https://news.ycombinator.com/item?id=49940394"
}
```

Comment (items mode):

```json
{
  "type": "comment", "id": 9272, "author": "dhouston", "createdAt": "2007-04-05T16:47:01Z",
  "storyId": 8863, "storyTitle": "My YC app: Dropbox - Throw away your USB drive",
  "parentId": 9224, "depth": 2, "rankAmongSiblings": 1, "replyCount": 1,
  "text": "1. re: the first part, many people want something plug and play. …",
  "links": [], "hnUrl": "https://news.ycombinator.com/item?id=9272", "commentSource": "algolia"
}
```

Job post (Who is hiring mode):

```json
{
  "type": "job-post", "id": 49922584, "month": "2026-10", "position": 2,
  "threadTitle": "Ask HN: Who is hiring? (October 2026)",
  "headerLine": "PrairieLearn (Remote US) — Full-Stack Software Engineer — TypeScript / Postgres / React / AI",
  "company": "PrairieLearn", "companyUrl": "https://www.prairielearn.com",
  "roles": ["Full-Stack Software Engineer", "AI"], "locations": ["US"],
  "remote": true, "workArrangement": ["remote"], "employmentType": ["full-time"],
  "salaryText": "$100k-$180k", "salaryMin": 100000, "salaryMax": 180000, "salaryCurrency": "USD", "salaryPeriod": "year",
  "otherHeaderParts": ["TypeScript", "Postgres", "React"],
  "applyUrls": ["https://www.prairielearn.com/jobs-ashby?..."],
  "text": "PrairieLearn (Remote US) — Full-Stack Software Engineer — … (full post)",
  "hnUrl": "https://news.ycombinator.com/item?id=49922584"
}
```

- Search results also have `matchedQuery`, `tags` and, for comments, `storyId`, `storyTitle`, `storyUrl`.
- Errors: `"type": "error"` with `input` and `error`. Not charged.
- The key-value store record `OUTPUT` holds the run summary (rows by type, API requests, stop reason).

### Pricing

Pay per event. You pay only for rows saved.

| Event | Price |
|---|---|
| Hacker News item (story, comment or job post) | **$0.0002** ($0.20 / 1,000) |
| Actor start | $0.00005 |

Compared with other Store Actors (public Store prices, 2026-10-01):

| Actor | Price per 1,000 items | Search / monitoring | Comment trees | Who is hiring parser |
|---|---|---|---|---|
| **Hacker News Scraper (this Actor)** | **$0.20** | yes, only-new | yes, flattened with depth | yes |
| gentle_cloud/hacker-news-scraper | $0.20 | – | – | – |
| ryanclinton/hackernews-search | $5.00 | yes | – | – |

Set **Maximum cost per run** in the run options to cap spending. The Actor stops cleanly at the limit, and in only-new mode the rows it could not save come back on the next run.

### FAQ

**Where does the data come from?**
From the official Hacker News API (`hacker-news.firebaseio.com`, live) and Algolia's public HN Search API (`hn.algolia.com`), which Hacker News links to for search. No HTML is scraped.

**How does only-new mode work?**
Returned IDs are stored per query / list / thread in a named key-value store derived from your input (`hacker-news-data-…`), or from `stateKey` if you set one. The first run returns the current results; later runs return only new ones. For a search sorted by date the Actor stops paging as soon as it reaches results from an earlier run, so scheduled runs are cheap.

**Why does a search return fewer results than Algolia's `nbHits`?**
Algolia matches typos and prefixes. With `exactMatch` on (default) typo tolerance is off and each word must appear at the start of a word, so `rust` no longer matches `trust`, but `postgres` still finds `PostgreSQL`. Turn it off to get Algolia's raw matching. Dropped results are free.

**Are comment trees complete?**
`auto` reads the whole tree from Algolia in one request (fast; Algolia can lag a few minutes behind HN on brand-new comments, and lists replies by time). When the Algolia tree is clearly incomplete, the Actor switches to the official API. Choose `firebase` for a live, exact HN ordering at every level. Deleted and flagged comments are skipped (their replies are kept) unless `includeDeletedComments` is on.

**How accurate is the Who is hiring parser?**
It is best-effort. On the October 2026 thread, `company` was found for 95% of 183 posts, location for 85%, remote/onsite for 93%, roles for 71% and salary text for 42% (numeric min/max for 37%; many posts do not state a salary). `headerStructured: false` marks posts that do not follow the `A | B | C` format. The full text is always included, so you can run your own parser or an LLM on it.

**Does it include personal data?**
Rows contain public HN usernames, which are part of the public data. Job posts may contain contact emails that companies posted for applications. Use the data in line with the GDPR and HN's guidelines; do not use it for spam.

### Use with AI agents / Apify MCP

- **Apify MCP server:** add `gazidev/hacker-news-data` to your MCP client (Claude Desktop, Cursor, VS Code) via `https://mcp.apify.com?actors=gazidev/hacker-news-data`. An agent can call it with `{"mode":"search","searchQueries":["your product"],"createdAfter":"7 days"}` and read what HN says about it.
- **API:** `POST https://api.apify.com/v2/acts/gazidev~hacker-news-data/run-sync-get-dataset-items?token=...` with the input JSON returns the rows directly, handy as a "Hacker News tool" for LangChain or LlamaIndex agents.
- **Scheduled alerts:** create a Task with `onlyNew: true`, schedule it, and connect the Slack, email, Google Sheets or webhook integration.

### More from the same developer

- [RSS Feed Reader](https://apify.com/gazidev/rss-feed-reader): any RSS, Atom or JSON Feed to JSON, only-new items.
- [Website to Markdown](https://apify.com/gazidev/website-to-markdown): any URL or whole site to clean Markdown for RAG.
- [Remote Jobs Aggregator](https://apify.com/gazidev/remote-jobs-aggregator): remote job boards in one dataset.

# Actor input Schema

## `mode` (type: `string`):

**Story lists**: front page (top), new, best, Ask HN, Show HN or job stories. **Search & monitor**: keyword or brand search in stories and/or comments (Algolia HN Search), with date range and points/comments thresholds. **Items + comments**: one or more stories (IDs or URLs) with the full comment tree. **Who is hiring**: the monthly 'Ask HN: Who is hiring?' thread split into job posts.

## `lists` (type: `array`):

Used in **Story lists** mode. `top` = the front page ranking.

## `maxItemsPerList` (type: `integer`):

Top N of each list, in HN ranking order (`rank`). HN keeps up to 500 (top/new/best) or 200 (ask/show/job).

## `searchQueries` (type: `array`):

Used in **Search & monitor** mode. One query per line, e.g. a brand, product, domain or topic. Use quotes for a phrase (`"web scraping"`). Each result has `matchedQuery`.

## `searchTags` (type: `string`):

Which kind of items to search.

## `sortBy` (type: `string`):

`date` = newest first, no limit on the number of results (use this for monitoring). `relevance` = Algolia ranking (points, matches), max 1,000 results per query.

## `createdAfter` (type: `string`):

Only items created on or after this date. Absolute (`2026-09-01`) or relative (`7 days`, `24 hours`).

## `createdBefore` (type: `string`):

Only items created before this date.

## `minPoints` (type: `integer`):

Stories only (comments have no public points). 0 = off.

## `minComments` (type: `integer`):

Stories only. 0 = off.

## `author` (type: `string`):

Optional. Only items by this user. Can be used without a query to get a user's latest stories or comments.

## `maxResultsPerQuery` (type: `integer`):

Newest first when sorted by date.

## `exactMatch` (type: `boolean`):

Algolia matches typos by default (`rust` also finds `trust`). On: typo tolerance is switched off and every query word must appear at the start of a word in the title, text, URL or story title (`postgres` still finds `PostgreSQL`). Dropped results are free.

## `itemIds` (type: `array`):

Used in **Items + comments** mode. HN item IDs (`8863`) or URLs (`https://news.ycombinator.com/item?id=8863`). Returns the item and all its comments.

## `includeComments` (type: `boolean`):

In **Story lists** and **Search** mode, also save the full comment tree of each story (each comment is a row and is charged as an item). Always on in Items mode.

## `maxCommentDepth` (type: `integer`):

1 = top-level comments only, 2 = plus direct replies, 0 = all levels.

## `maxCommentsPerItem` (type: `integer`):

0 = all. Comments are taken in thread order (top-level by HN ranking, each followed by its replies).

## `includeDeletedComments` (type: `boolean`):

Off: deleted and flagged (dead) comments are skipped, their replies are kept.

## `commentSource` (type: `string`):

`auto`: whole tree from Algolia in one request (fast, top level in HN ranking order, replies by time), falls back to the official API when Algolia is incomplete. `firebase`: official API only (live, exact HN order, one request per comment - slower on big threads).

## `hiringMonth` (type: `string`):

Used in **Who is hiring** mode. `latest`, `2026-09` or `September 2026`.

## `hiringKeywords` (type: `array`):

Optional. Keep only job posts containing these words (whole words, case-insensitive), e.g. `python`, `rust`, `Berlin`. Filtered posts are free.

## `hiringKeywordsMode` (type: `string`):

Any keyword or all keywords.

## `remoteOnly` (type: `boolean`):

Keep only posts detected as remote (`remote: true`).

## `onlyNew` (type: `boolean`):

Remembers what was already returned (in a named key-value store derived from this input) and outputs only new stories, search results, comments or job posts. Schedule the Actor hourly or daily for brand alerts, new comments on a thread or new job posts. The first run returns the current results.

## `stateKey` (type: `string`):

Optional. By default the memory is tied to this input. Set a name to keep the memory when you edit the input.

## `maxItems` (type: `integer`):

Stop after this many rows (stories + comments + job posts). 0 = no limit. You pay only for rows saved.

## `includeHtml` (type: `boolean`):

Add `textHtml` with HN's HTML next to the plain `text`.

## `maxConcurrency` (type: `integer`):

Parallel requests to the official API (Algolia is always limited to 4).

## `requestTimeoutSecs` (type: `integer`):

Failed requests are retried 3 times with backoff.

## Actor input object example

```json
{
  "mode": "lists",
  "lists": [
    "top"
  ],
  "maxItemsPerList": 30,
  "searchQueries": [
    "Apify",
    "\"web scraping\""
  ],
  "searchTags": "story",
  "sortBy": "date",
  "minPoints": 0,
  "minComments": 0,
  "maxResultsPerQuery": 50,
  "exactMatch": true,
  "itemIds": [
    "https://news.ycombinator.com/item?id=8863"
  ],
  "includeComments": false,
  "maxCommentDepth": 0,
  "maxCommentsPerItem": 0,
  "includeDeletedComments": false,
  "commentSource": "auto",
  "hiringMonth": "latest",
  "hiringKeywordsMode": "any",
  "remoteOnly": false,
  "onlyNew": false,
  "maxItems": 0,
  "includeHtml": false,
  "maxConcurrency": 20,
  "requestTimeoutSecs": 30
}
```

# Actor output Schema

## `stories` (type: `string`):

No description

## `comments` (type: `string`):

No description

## `jobs` (type: `string`):

No description

## `errors` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "lists",
    "lists": [
        "top"
    ],
    "maxItemsPerList": 30,
    "searchQueries": [
        "Apify",
        "\"web scraping\""
    ],
    "maxResultsPerQuery": 50,
    "itemIds": [
        "https://news.ycombinator.com/item?id=8863"
    ],
    "hiringMonth": "latest"
};

// Run the Actor and wait for it to finish
const run = await client.actor("gazidev/hacker-news-data").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "lists",
    "lists": ["top"],
    "maxItemsPerList": 30,
    "searchQueries": [
        "Apify",
        "\"web scraping\"",
    ],
    "maxResultsPerQuery": 50,
    "itemIds": ["https://news.ycombinator.com/item?id=8863"],
    "hiringMonth": "latest",
}

# Run the Actor and wait for it to finish
run = client.actor("gazidev/hacker-news-data").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "lists",
  "lists": [
    "top"
  ],
  "maxItemsPerList": 30,
  "searchQueries": [
    "Apify",
    "\\"web scraping\\""
  ],
  "maxResultsPerQuery": 50,
  "itemIds": [
    "https://news.ycombinator.com/item?id=8863"
  ],
  "hiringMonth": "latest"
}' |
apify call gazidev/hacker-news-data --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gazidev/hacker-news-data"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/yuxUbbq8cmCpfNYzi/builds/6twiFfKaZJ6Q0DBuw/openapi.json
