# Hacker News Scraper — Stories, Comments & Who Is Hiring (`keyman98/hacker-news-live-api`) Actor

Get Hacker News stories from the official HN API: top, new, best, Ask HN, Show HN and jobs. Filter by keyword and minimum score, optionally with comments. Clean text, direct links, JSON/CSV/Excel.

- **URL**: https://apify.com/keyman98/hacker-news-live-api.md
- **Developed by:** [KeyMan98](https://apify.com/keyman98) (community)
- **Categories:** News, Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.70 / 1,000 story scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Hacker News Scraper — Stories, Comments & Who Is Hiring

Get live Hacker News stories — **Top**, **New**, **Best**, **Ask HN**, **Show HN**, and the monthly **Who Is Hiring** thread — with comments, filtered by keyword or score. Pulls data through Hacker News' own official, public **Firebase API** (`hacker-news.firebaseio.com`) — no HTML scraping of news.ycombinator.com, no third-party data provider — and exports the results to CSV, Excel, or JSON.

### What you get (output fields)

For each story, one dataset row with:

- `id` / `type` — the Hacker News item ID, and its type as reported by the API (e.g. `story`, `job`).
- `title` / `url` — the story's title and, if it links out, its external URL.
- `text` — for text-only posts (most Ask HN, Show HN, and job posts): the body as clean plain text — HTML entities decoded, paragraphs turned into line breaks, links kept as plain URLs.
- `by` — the author's public Hacker News username (public data from the official API).
- `score` — points at the time of the run.
- `descendants` — total comment count on the story, as reported by HN (independent of how many were actually fetched).
- `time` — when the story was submitted (ISO 8601, UTC).
- `hnUrl` — the story's discussion page on news.ycombinator.com.
- `domain` — the domain of `url`, without `www.` (null for text-only posts).
- `comments` — a list of comments (`id`, `by`, `text`, `time`, `parent`, `depth`), only if "Include comments" was on; otherwise null.
- `error` — set only on rows that could not be resolved (see below); null otherwise.

### Who it's for

- **Tech trend / product monitoring** — track what's on the Hacker News front page for a brand, technology, or company name.
- **Research and datasets** — pull structured HN discussions (stories + comment trees) for analysis or as training/evaluation data.
- **"Who is Hiring" tracking** — scan the monthly Ask HN hiring thread for a role, technology, or "remote" without reading it by hand.
- **Alerts** — run on a schedule and filter by keyword/score to get notified only when something relevant appears.

This Actor does not do full-text search over HN's history — see "Limitations" below.

### How to use

1. **Story type** — `top`, `new`, `best`, `ask` (Ask HN), `show` (Show HN), or `job` (Who is Hiring / job posts). Ignored if "Item IDs" is set.
2. **Max items** — stop after this many stories pass the filters below.
3. **Keywords (optional)** — keep only stories whose title or text contains at least one of these words (case-insensitive).
4. **Minimum score (optional)** — keep only stories with at least this many points.
5. **Include comments** — fetch each story's comment tree too (breadth-first, shallow replies first).
6. **Max comments per story** — cap on how many comments to fetch per story, if comments are included.
7. **Item IDs (optional)** — specific HN item IDs to fetch directly instead of a story list; overrides "Story type".

### Input example (JSON)

```json
{
  "storyType": "top",
  "maxItems": 30,
  "keywords": [],
  "minScore": null,
  "includeComments": false,
  "maxCommentsPerStory": 20,
  "itemIds": []
}
```

### Output example (JSON)

```json
{
  "id": 1,
  "type": "story",
  "title": "Y Combinator",
  "url": "http://ycombinator.com",
  "text": null,
  "by": "pg",
  "score": 61,
  "descendants": 19,
  "time": "2006-10-09T18:21:51+00:00",
  "hnUrl": "https://news.ycombinator.com/item?id=1",
  "domain": "ycombinator.com",
  "comments": null,
  "error": null
}
```

### If an item ID is not found or can't be fetched

That entry becomes one **error row**: `error` is set to a short explanation, every other field is null. The run does not fail, the rest of the list keeps running, and **you are not charged** for that row. A story that's been deleted or marked dead on Hacker News is instead skipped entirely — no row at all, since that's HN's own removal, not a fetch problem.

### Pricing

Pay only for stories actually returned, after filters — nothing charged for an item ID that could not be resolved, and nothing extra for comments (they're included in the story's own charge). Pricing model: **pay-per-event**.

| Event | When it's charged | Price |
| --- | --- | --- |
| `item-scraped` | a story was returned in the results (after filters) | 0.0007 USD |

### Limitations

- **No full-text search over HN's history.** This Actor only reads live lists (top/new/best/ask/show/job, each up to 500 items) and specific item IDs you already know. For searching HN's entire archive by keyword, use the separate, community-run **Algolia HN Search API** (`hn.algolia.com/api`) — not used here.
- **No per-user data.** Only the public `by` username on each story/comment is returned — no karma, no submission history.
- Deleted or dead stories are skipped, never returned as a row.
- `comments` reflects what was actually fetched (bounded by "Max comments per story"); `descendants` is HN's own total comment count, and the two can differ on a busy story.
- Job posts (`storyType` `job`) have no `score` field — setting "Minimum score" always excludes them, regardless of the value.

### FAQ

#### Am I charged if an item ID is not found?

No. You are only charged for stories actually returned in the results, after filters.

#### Where does the data come from?

The official, public Hacker News Firebase API (`hacker-news.firebaseio.com/v0`), documented at `github.com/HackerNews/API`. No HTML page of news.ycombinator.com is ever fetched or parsed.

#### Does fetching comments cost extra?

No. Comments are included in the same per-story charge — only whether "Include comments" is on changes what's inside the `comments` field, not the price.

#### How do keyword and score filters work?

`keywords` matches as a whole word, case-insensitively, against the story's title and text — kept if at least one keyword appears in either. "AI" won't match inside "said" or "email", and keywords like "C++" or ".NET" still match correctly at the edge of a sentence. `minScore` keeps only stories with at least that many points — note that job posts (`storyType` `job`) have no score at all, so setting `minScore` always excludes them. Both filters apply before anything is added to the dataset, so filtered-out stories are never charged.

#### How often is the data updated?

Every run re-fetches live from the HN API — results reflect the score/comment count at run time, not a cached snapshot.

#### Can I monitor Hacker News on a schedule?

Yes. Set "Story type" and a keyword or score filter, then run this Actor on a schedule (Apify's built-in scheduler).

#### Can I re-check specific stories later?

Yes — pass their IDs in "Item IDs" (found in a story's `hnUrl`, after `?id=`). This ignores "Story type" and fetches each ID once, even if listed twice.

#### Can I use this through the Apify API or an MCP server?

Yes, like any Apify Actor — through the standard Apify API, or through the Apify MCP server if you use Claude, Cursor, or another MCP-enabled client.

### Export

Results can be downloaded from the Apify dataset as JSON, CSV, or Excel, or accessed via the Apify API.

# Actor input Schema

## `storyType` (type: `string`):

Which Hacker News list to pull stories from. Ignored if "Item IDs" is set.

## `maxItems` (type: `integer`):

Stop after this many stories pass the filters below (keywords/minimum score).

## `keywords` (type: `array`):

Keep only stories whose title or text contains at least one of these words as a whole word (case-insensitive, e.g. "AI" does not match "said" or "email"; "C++" and ".NET" also match correctly). Leave empty to keep every story.

## `minScore` (type: `integer`):

Keep only stories with at least this many points. Leave empty for no minimum. Note: Hacker News job posts (storyType "job") have no score at all, so they are always excluded once this is set.

## `includeComments` (type: `boolean`):

Fetch each story's comments too. Adds to run time, not to the price - comments are included in the same per-story charge.

## `maxCommentsPerStory` (type: `integer`):

Stop fetching a story's comments after this many (breadth-first: shallow replies before deeper ones). Only used if "Include comments" is on.

## `itemIds` (type: `array`):

Specific Hacker News item IDs to fetch instead of a story list (e.g. to re-check known stories). If set, "Story type" is ignored.

## Actor input object example

```json
{
  "storyType": "top",
  "maxItems": 30,
  "keywords": [],
  "includeComments": false,
  "maxCommentsPerStory": 20,
  "itemIds": []
}
```

# Actor output Schema

## `results` (type: `string`):

All results in the default dataset (JSON, CSV, Excel).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "storyType": "top",
    "keywords": [],
    "itemIds": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("keyman98/hacker-news-live-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "storyType": "top",
    "keywords": [],
    "itemIds": [],
}

# Run the Actor and wait for it to finish
run = client.actor("keyman98/hacker-news-live-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "storyType": "top",
  "keywords": [],
  "itemIds": []
}' |
apify call keyman98/hacker-news-live-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,keyman98/hacker-news-live-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/FR9m4N2QlJCaP5ZT7/builds/XgUYJkfTzG6VsZl3f/openapi.json
