Hacker News Scraper: Stories, Comments, Search & Alerts
Pricing
from $0.50 / 1,000 stories
Hacker News Scraper: Stories, Comments, Search & Alerts
Hacker News front page, new, best, Ask HN, Show HN and job lists, or keyword search with date, points and comment filters. Optional comments, flattened with depth. Schedule it with "Only new items" for brand and topic alerts. Official HN and Algolia APIs. USD 0.50 per 1,000 stories.
Pricing
from $0.50 / 1,000 stories
Rating
0.0
(0)
Developer
JT Palms
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
an hour ago
Last modified
Categories
Share
Get Hacker News as clean JSON, CSV or Excel. Pick the lists you want (front page, newest, best, Ask HN, Show HN, jobs) or search the whole HN archive by keyword, with date, points and comment filters. Add each story's comments, flattened into rows with their depth. Turn on Only new items and schedule it to get an alert whenever a new story mentions your brand, product or topic.
It uses the two official ways to read Hacker News: the Hacker News API that Y Combinator publishes, and the Algolia HN Search API. No login, no page scraping, no browser.
USD 0.50 per 1,000 stories, USD 0.20 per 1,000 comments. Failed lists or searches are free, and monitor runs with nothing new cost nothing.
What people use it for
- Brand and topic monitoring to Slack. Search for your company, product or competitors, turn on Only new items, and schedule it every hour. Connect the task to Slack, email, Google Sheets or a webhook and each new mention lands there. Search Stories and comments to catch mentions deep in threads too.
- Launch tracking. Follow Show HN and Launch HN posts, or watch the front page for stories from your domain. See points and comment counts climb, and pull the comments to read the feedback.
- Research datasets. Pull every story about a topic for a date range (
"rust", 2024-01-01 to 2024-12-31, at least 50 points) for trend analysis, or a front-page snapshot every day for a longitudinal dataset. - AI agents and LLM pipelines. Comments come as plain text with story, parent and depth, ready for summarising a discussion, sentiment analysis or RAG. Agents can call it through the Apify API or MCP server to answer "what does HN think about X?".
- Hiring and job signals. The Jobs list gives YC company job posts; keyword search on jobs finds roles by technology.
Sample output
Real rows from the front page with comments, trimmed:
[{"type": "story","kind": "story","id": 49880036,"title": "Pirating the Pirates","url": "https://mubi.com/en/notebook/posts/pirating-the-pirates","domain": "mubi.com","author": "piotrgrabowski","points": 205,"commentCount": 67,"createdAt": "2026-09-28T15:54:15.000Z","hnUrl": "https://news.ycombinator.com/item?id=49880036","text": null,"storyId": 49880036,"depth": 0,"rank": 1,"source": "top"},{"type": "comment","kind": "comment","id": 49881259,"author": "schlauerfox","createdAt": "2026-09-28T17:17:00.000Z","hnUrl": "https://news.ycombinator.com/item?id=49881259","text": "One thing to note, is that the library of congress has the power to create the exceptions to the DMCA. The EFF lobbies for expansion of the exact powers the article is advocating for.\nhttps://www.eff.org/issues/dmca-rulemaking","storyId": 49880036,"storyTitle": "Pirating the Pirates","parentId": 49880036,"depth": 1,"source": "top"}]
A keyword search result (postgres, stories and comments):
{"type": "story","kind": "show","id": 49882674,"title": "Show HN: iCli - Postgres health reporter and index optimizer","url": "https://github.com/abhiraj-ku/iCli","domain": "github.com","author": "abhirajabhi312","points": 2,"commentCount": 0,"createdAt": "2026-09-28T18:51:27.000Z","hnUrl": "https://news.ycombinator.com/item?id=49882674","source": "search: postgres"}
Every row has the same fields, so CSV and Excel exports line up:
| Field | What it is |
|---|---|
type | story, job, poll or comment. |
kind | story, ask, show, launch, tell, job, poll or comment. Text posts without a link count as ask, as on HN. |
id, hnUrl | The HN item ID and its page on news.ycombinator.com. |
title, url, domain | Story title, the link it points to, and that link's domain without www.. |
author | The HN username that posted it. |
points, commentCount | Upvotes and total comments at the time of the run. HN does not publish comment scores. |
createdAt | When it was posted, ISO 8601 in UTC. |
text | Plain text of an Ask HN, Show HN or job post, or of a comment. HTML is removed, links are written out in full, paragraphs become blank lines. |
storyId, storyTitle, parentId, depth | For comments: the story, the item it replies to, and how deep it is (1 = reply to the story). Stories have depth 0. |
rank | Position in the list (1 = top of the front page). Lists only. |
source | Which list or search keyword produced the row. |
How to use it
Lists
- Leave What to get on Lists and pick one or more lists. Top is the front page ranking.
- Optional: Minimum points, Minimum comments, Posted after and Must contain (any of) (for example
rust,postgres*). - Set Max stories per list or keyword (default 100).
- Click Start. Export as JSON, CSV or Excel, or read the dataset through the Apify API.
Keyword search
- Set What to get to Keyword search and add one keyword or phrase per line. Quotes make an exact phrase:
"open source". - Choose Search in: stories, comments, both, or only Ask HN, Show HN, Launch HN, jobs, polls or the current front page. Ask HN also covers Tell HN and other text posts; the
kindfield tells them apart. - Optional filters: Posted after and Posted before (a date like
2026-01-01or a relative value like7 days), Minimum points, Minimum comments. - Sort newest first (any number of results) or by relevance (up to 1,000 per keyword).
Comments
Turn on Include comments. Each story is followed by its comments, in the order HN ranks them. Comment depth 1 gives only top-level comments; 2 adds the replies to them. Max comments per story caps the count, taking top-level comments first, so a small number gives you the best of the discussion.
Set it up as an alert
- Turn on Only new items.
- Choose what the first run does: All current items returns what matches now (for a search without Posted after, the last 7 days); None, just set the starting point returns nothing and only records what exists, so you hear only about new items.
- Save it as a task and add a schedule, for example every hour.
- In the task's Integrations tab, connect Slack, email, Google Sheets, Zapier, Make or a webhook.
Examples:
- New stories mentioning your product: Keyword search,
yourproduct, Search in Stories and comments, Only new items, hourly. - Front page stories about your field: Lists, Top, Must contain
robotics, Only new items, every 30 minutes. - Big launches only: Keyword search with no keyword, Search in Show HN, Minimum points
100, Only new items, daily.
With Minimum points the alert fires when a story reaches the threshold, not when it is posted, so a slow climber still triggers once. If you run two schedules on the same list or keyword with different filters, give each a Monitor name (under Monitoring) so they keep separate memories.
Pricing
| What | Price |
|---|---|
| Story, Ask HN, Show HN, job or poll in the results | USD 0.0005 (USD 0.50 per 1,000) |
| Comment in the results | USD 0.0002 (USD 0.20 per 1,000) |
| A list or search that fails | Free |
| Monitor run where nothing is new | Free |
Examples: the front page (30 stories) with 20 comments each is 30 x 0.0005 + 600 x 0.0002 = USD 0.135. A brand alert that finds 5 new mentions a day costs about USD 0.08 a month. 10,000 stories for a research dataset cost USD 5.
Set a maximum cost per run in the run options. The actor stops cleanly when it gets there, and with Only new items anything it did not save is kept for the next run.
Limits
- Lists hold what HN holds: Top and New up to 500 stories, Best 200, Ask, Show and Jobs up to 200.
- Relevance search stops at 1,000 results per keyword (a limit of the search API). Newest-first search has no such limit; the actor walks back through time.
- The search index can be a few minutes behind the live site, and its points and comment counts are refreshed periodically. With Include comments, points and comment counts are re-read live from the official API.
- Only new items remembers up to 20,000 item IDs per list or keyword in your
hacker-news-monitorkey-value store. A search monitor looks back 48 hours before the newest item it has seen, so a story that crosses Minimum points more than two days after it was posted is not reported. - Deleted and flagged ("dead") items are skipped. The replies under a deleted comment are kept; the replies under a flagged one are not, as on HN by default.
- The actor does not log in, vote, post, or read user profiles.
FAQ
Is this allowed? Yes. It only uses the official Hacker News API, which Y Combinator publishes for exactly this kind of use, and the public Algolia HN Search API that powers the search box at the bottom of every HN page. It never scrapes news.ycombinator.com pages. The content belongs to its authors and to Hacker News: quote and link back, and check HN's terms before republishing large amounts of it.
Does it collect personal data? It returns what HN shows publicly on each post: the text and the author's username. It does not read user profiles, "about" fields, karma or submission histories.
Why did the first alert run return items? With All current items the first run returns what matches now. Choose None, just set the starting point to hear only about items posted after the first run.
Can I start over? Delete the matching record in the hacker-news-monitor key-value store (Storage, Key-value stores), or use a new Monitor name.
I get HTTP 429 from search. The search API allows 10,000 requests an hour per IP address, and cloud IPs are shared. Turn on a proxy under Advanced.
Something missing or wrong? Open an issue with your input.