# Hacker News Search Scraper (`67-labs/hacker-news-search-scraper`) Actor

Search Hacker News stories and comments by keyword, date and points. Title, link, author, points, comments, text. CSV or Sheets.

- **URL**: https://apify.com/67-labs/hacker-news-search-scraper.md
- **Developed by:** [Mokksh Bhatt](https://apify.com/67-labs) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 result (one story or comment)s

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Hacker News Search Scraper — find stories and comments by keyword, date and points

Search Hacker News for any word or phrase and get the stories and comments as a clean table: title, link, site, author, points, number of comments, date, the full text and the Hacker News link. Filter by type (stories, comments, Ask HN, Show HN), by date range and by minimum points. Export to CSV, JSON, Excel or Google Sheets. It reads Hacker News through its open search service, so it is fast and stable.

**Who uses this:** founders checking what people say about their product or a competitor, marketers and PR teams watching mentions, developer relations teams, investors and analysts spotting trends, researchers who need discussion data, AI agents that need current opinions on a topic.

### What you get per result

| Field | What it is |
|---|---|
| `id`, `hnUrl` | The item's ID and its Hacker News link |
| `type` | `story`, `show_hn`, `ask_hn`, `comment` or `poll` |
| `title`, `url`, `domain` | Title of the story, the link it points to, and that link's site (empty for comments and text posts) |
| `author` | The Hacker News username |
| `points`, `numComments` | Points, and number of comments on a story |
| `createdAt` | When it was posted (ISO time, UTC) |
| `text` | The text of the post or comment, cleaned from HTML |
| `storyId`, `storyTitle`, `parentId` | For a comment: the story it belongs to and the item it answers |
| `query` | Which of your search texts found it |
| `scrapedAt` | When the row was collected |

### Price

**$0.01 per run + $1 per 1,000 results.** That is two pay-per-event charges: `actor-start` ($0.01) once per run and `result` ($0.001) per story or comment. No compute fees on top. The free Apify plan ($5 per month credit) covers about 4,900 results a month at no cost. A run that the search service blocks before it returns any data is not charged.

### How to use

1. Open the Input tab and type your search texts, one per line, for example `web scraping` or `YC W26`.
2. Pick the **Type**, and optionally a date range and a minimum of points. Set **Max results per search**. This is also your cost cap.
3. Click Start. When it is done, open the Output tab and export the table as CSV, JSON or Excel, or send it to Google Sheets with Apify's Google Sheets integration.

### Input examples

Newest Show HN posts about scraping, with at least 10 points:

```json
{ "queries": ["scraping"], "type": "show_hn", "minPoints": 10, "maxResults": 100 }
```

Every comment that mentions your product this month:

```json
{
    "queries": ["apify", "crawlee"],
    "type": "comment",
    "dateFrom": "2026-09-01",
    "dateTo": "2026-09-30",
    "maxResults": 500
}
```

The best-matching stories about a topic:

```json
{ "queries": ["vector database"], "type": "story", "sort": "relevance", "maxResults": 50 }
```

### Output example

```json
{
    "id": "49320034",
    "type": "show_hn",
    "title": "Show HN: PageSieve, a web scraping browser extension",
    "url": null,
    "domain": null,
    "author": "kajm",
    "points": 17,
    "numComments": 1,
    "createdAt": "2026-08-16T13:45:27Z",
    "text": "PageSieve[1] is a browser extension for scraping data from different websites from within your browser. Currently it's  Firefox only but I am open to porting it to other browsers ...",
    "storyId": "49320034",
    "storyTitle": "Show HN: PageSieve, a web scraping browser extension",
    "parentId": null,
    "hnUrl": "https://news.ycombinator.com/item?id=49320034",
    "query": "web scraping",
    "scrapedAt": "2026-09-26T10:00:00.000Z"
}
```

The text is shortened in this example.

### Limits

- Hacker News search returns at most **1,000 results for one search text**. Use a date range or a higher minimum of points to reach more, or split the search into smaller periods.
- Public data only: everything on Hacker News is public.
- Search matches words in the title, the text and the link. It is not an exact-phrase search unless you put the phrase in quotes.
- When several items share the exact same second, the search service can repeat one across pages. The actor removes such repeats, so you are not charged twice, and a very rare item can be missing.

### FAQ

**Can I watch mentions every day?** Yes. Schedule the actor in Apify with `sort` set to `date`, a `dateFrom` of yesterday, and compare `id` values between runs.

**Can I search comments only?** Yes. Set **Type** to Comments.

**Can an AI agent use it?** Yes. It works through Apify's API and MCP server. Give it search texts and a `maxResults` limit.

**Something is wrong. What now?** Open an issue on the Issues tab with the search and the input you used. Field requests are welcome.

### Changelog

- **0.1** First release: search by keyword, type, date and points; cleaned text; pay per event.

# Actor input Schema

## `queries` (type: `array`):

One entry per search: a word or phrase, for example web scraping or "YC W26". Each search is run on its own.

## `type` (type: `string`):

What to search: everything, stories, comments, Ask HN posts or Show HN posts.

## `sort` (type: `string`):

Newest first, or the best match first.

## `dateFrom` (type: `string`):

Only items posted on or after this day.

## `dateTo` (type: `string`):

Only items posted on or before this day.

## `minPoints` (type: `integer`):

Only items with at least this many points. Use 10 or more to skip posts nobody noticed.

## `maxResults` (type: `integer`):

Stop after this many results for each search text. The most Hacker News search can return for one search is 1,000. You pay per result, so this is also your cost cap.

## `proxyConfiguration` (type: `object`):

Leave off. The search service is open.

## Actor input object example

```json
{
  "queries": [
    "web scraping",
    "apify",
    "\"show hn\" scraper"
  ],
  "type": "any",
  "sort": "date",
  "minPoints": 0,
  "maxResults": 20,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

One row per story or comment: title, link, author, points, comments, text, Hacker News link.

## `resultsCsv` (type: `string`):

Same rows as CSV for Google Sheets or Excel.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "web scraping"
    ],
    "maxResults": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("67-labs/hacker-news-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["web scraping"],
    "maxResults": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("67-labs/hacker-news-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "web scraping"
  ],
  "maxResults": 20
}' |
apify call 67-labs/hacker-news-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,67-labs/hacker-news-search-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/rHvntdFg6XZJrfrRX/builds/cj92z3YNdg5OelkVU/openapi.json
