# Hacker News Scraper: Story & Comment Search (`recordsdata/hackernews-search-scraper`) Actor

Scrape Hacker News: full-text search across 40M+ stories and comments with points, dates and type filters (Ask HN, Show HN, jobs, front page), plus user profiles with karma. Export CSV, Excel, JSON, XML.

- **URL**: https://apify.com/recordsdata/hackernews-search-scraper.md
- **Developed by:** [RecordsData](https://apify.com/recordsdata) (community)
- **Categories:** News, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

<p align="center">
  <img src="https://api.apify.com/v2/key-value-stores/AAm3a1h3Z9nYfrvh9/records/banner?v=3" alt="RecordsData" width="100%" />
</p>

## 🟠 Hacker News Scraper - Story & Comment Search - RecordsData

> 🚀 **The Hacker News Scraper is an Apify actor that searches 40M+ Hacker News stories and comments by keyword, with points, comment counts, dates and direct links.** Filter by type (stories, Ask HN, Show HN, jobs, the live front page), minimum points and date range; sort by relevance or newest; and pull user profiles with karma. A search for "web scraping" returns 178 stories, the top one at 1,057 points, measured live. Export to CSV, Excel, JSON or XML.

HN is where developer sentiment forms; this actor turns any topic into a ranked, dated dataset of what that community said about it.

| 🎯 Target Audience | 💡 Primary Use Cases |
|---|---|
| DevTool marketers and founders | Track every mention of your product or category |
| Investors and analysts | Developer sentiment on technologies over time |
| Recruiters | Who is hiring / who wants to be hired datasets |
| Researchers | Tech discourse corpora with engagement signals |

### 📋 What the Hacker News Scraper does

- **Story search**: title, external URL, author, points, comment count, type, self-text and exact HN link for any query, deep-paginated.
- **Type filters**: stories, Ask HN, Show HN, jobs, polls, or the current front page.
- **Signal filters**: minimum points and date range, so you only pay for stories that mattered.
- **Comment search**: full-text across comments with author, parent story and date.
- **User profiles**: karma, about and account age per username.

> 💡 **Why it matters:** an HN front page appearance moves markets and traffic; this actor makes the whole archive queryable as data.

### 📊 Output of the Hacker News export

Real sample from a live run:

```json
{
  "recordType": "story",
  "hnId": "27968148",
  "title": "Congrats! Web scraping is legal! (US precedent)",
  "url": "https://example-source.com/article",
  "author": "ehurynovich",
  "points": 1057,
  "commentCount": 388,
  "storyType": "story",
  "createdAt": "2021-07-26T12:31:11Z",
  "hnUrl": "https://news.ycombinator.com/item?id=27968148",
  "error": null
}
```

You can download the dataset in various formats such as JSON, CSV, Excel or XML.

### ✨ Why choose this Hacker News scraper

- **Stories + comments + users** in one actor; alternatives cover stories only.
- **Points and date filters** to slice signal from 18 years of archive.
- **Six post types** including the live front page as a monitorable list.
- **Official search API sourcing**: fast, complete, no HTML parsing.

### 📈 How this Hacker News scraper compares to alternatives

Measured against the HN actors on the Apify Store (September 2026):

| | This actor | The three alternatives |
|---|---|---|
| Comment search | Yes | No |
| User profiles | Yes | No |
| Min-points + date filters | Yes | Partial |
| Ask/Show/Jobs/front-page types | All | Usually one |
| Price per 1,000 stories | $5.50 | $0.20-$5.00 with fewer capabilities |

### 🚀 How to use the Hacker News Scraper

1. Create a free Apify account (with $5 of credit) at console.apify.com.
2. Open this actor's page and click **Try for free**.
3. Type a query and pick a story type and sort.
4. Optionally set minimum points, dates, comments and usernames.
5. Click **Start** and download CSV, Excel, JSON or XML.

### 💼 Business use cases

#### Brand monitoring

Every HN mention of your product, scored and dated.

#### Market research

What developers praised or trashed in any category, over years.

#### Hiring intelligence

Monthly Who-is-Hiring style datasets via job-type filters.

#### Content strategy

Which headlines about your topic reached the front page.

### 🔌 Automating the Hacker News Scraper

Connect to **Make**, **Zapier**, **Slack**, **Airbyte**, **GitHub** or **Google Drive**: hourly mention alerts to Slack, weekly sentiment exports, or archive syncs to your warehouse.

### 🤖 Ask an AI assistant about this scraper

> "I need every HN story about 'postgres' with 100+ points since 2023, plus matching comments, as CSV. Would the Hacker News Scraper on Apify (apify.com/recordsdata/hackernews-search-scraper) do this?"

### ❓ Frequently Asked Questions

#### 🟠 How far back does the archive go?

To Hacker News's start in 2007; date filters let you slice any window.

#### 🔎 What does relevance sort mean?

The search index's blend of text match and popularity; switch to "newest first" for monitoring workflows.

#### 💬 Do comment rows include their story?

Yes, each comment carries its parent story title and ID for joining.

#### ⚖️ Is this allowed?

The actor uses the official public HN Search API (by Algolia) within its generous limits.

#### 🆓 Can I try it for free?

Yes, free users get a 10-row preview. Paid Apify plans unlock full runs.

Found a bug or missing field? Open the **Issues tab** on this actor's page. Custom solutions available on request.

# Actor input Schema

## `includeStories` (type: `boolean`):

Scrape stories matching the query and filters.

## `searchQuery` (type: `string`):

Full-text search across titles and content. Leave empty for everything (use filters).

## `storyType` (type: `string`):

Which kind of posts to scrape.

## `sortBy` (type: `string`):

Relevance (best match, popularity-weighted) or date (newest first).

## `minPoints` (type: `integer`):

Only items with at least this many upvotes (0 = any).

## `dateFrom` (type: `string`):

Only items created after this date (YYYY-MM-DD).

## `dateTo` (type: `string`):

Only items created before this date (YYYY-MM-DD).

## `includeComments` (type: `boolean`):

Also search comments matching the query (one row per comment).

## `usernames` (type: `array`):

HN usernames for profile rows (karma, about, since).

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## Actor input object example

```json
{
  "includeStories": true,
  "searchQuery": "web scraping",
  "storyType": "story",
  "sortBy": "relevance",
  "minPoints": 0,
  "includeComments": false,
  "usernames": [],
  "maxItems": 10
}
```

# Actor output Schema

## `overview` (type: `string`):

Key fields per row

## `fullData` (type: `string`):

Complete dataset with all fields

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQuery": "web scraping",
    "dateFrom": "",
    "dateTo": "",
    "usernames": [],
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("recordsdata/hackernews-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQuery": "web scraping",
    "dateFrom": "",
    "dateTo": "",
    "usernames": [],
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("recordsdata/hackernews-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQuery": "web scraping",
  "dateFrom": "",
  "dateTo": "",
  "usernames": [],
  "maxItems": 10
}' |
apify call recordsdata/hackernews-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,recordsdata/hackernews-search-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ufJo5myr4rvxDWdfH/builds/4dMCYm2drhQkJgNnd/openapi.json
