Hacker News Scraper - Stories, Ask/Show HN, Comments avatar

Hacker News Scraper - Stories, Ask/Show HN, Comments

Pricing

from $0.97 / 1,000 items

Go to Apify Store
Hacker News Scraper - Stories, Ask/Show HN, Comments

Hacker News Scraper - Stories, Ask/Show HN, Comments

Search Hacker News stories, Show HN, Ask HN and comments, or pull the live front page. Each row has the title, link, author, points, comment count, the text and the HN thread link. Filter by minimum points or comments. No account or API key. $1.00 per 1,000 items.

Pricing

from $0.97 / 1,000 items

Rating

0.0

(0)

Developer

Dami's Studio

Dami's Studio

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

2

Monthly active users

a day ago

Last modified

Share

Scrape Hacker News stories, Show HN, Ask HN and comments: type a query, pick what kind of item you want, and the matches come back as rows with points, author, comment count, the story's own link and the HN thread link. Or pull whatever is sitting on the front page right now. No account, no API key.

One thing to know before the first run. The search box on the input form starts out filled in with openai, and the query is applied to the front page as well. Leave it there with Front page selected and you get the few front-page items that mention OpenAI, not the front page.

InputA search query and an item type
OutputOne row per story, post or comment: title, link, author, points, comment count, text, thread link
Ceiling1,000 items per run
Account neededNone, and no API key
Price$1.00 per 1,000 items, flat on every plan. The free plan's $5 a month covers about 5,000

๐Ÿ” What Hacker News Scraper does

Pick the item type and you get stories, Show HN posts, Ask HN posts, comments, or the current front page. Leave the query empty and you get the newest items of that type instead of a keyword search. Sort by relevance or by date. Set a minimum score, or a minimum number of comments, to drop the quiet stories before they reach your dataset.

It pages through results 50 at a time until it has what you asked for or the results run out. Items that appear on two pages are dropped once, so the same story cannot land twice in one run.

There is nothing to sign up for and no quota to watch.

๐Ÿ“‹ What data you get from each Hacker News item

What you getField
The item's id and what kind of item it isobjectId, type
The title and the link it points totitle, url
Who posted it and whenauthor, createdAt
Points and comment countpoints, numComments
The post or comment body, as plain texttext
The thread on news.ycombinator.comhnUrl

โ–ถ๏ธ How to scrape Hacker News

  1. Open Hacker News Scraper and click Try for free.
  2. Type your keywords into Search query, or clear the box to get the newest items instead.
  3. Pick an Item type. Stories is the default, comments is the one you want for mention tracking.
  4. Set Max items. Start around 20 to see the shape of the output, then click Start.
  5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.

๐Ÿ’ฐ How much does it cost to scrape Hacker News?

$1.00 per 1,000 items. Flat on every Apify plan, no volume tiers. On the free plan, the $5 Apify gives you each month covers about 5,000 items.

You pay per row delivered, so maxItems is your budget dial and the default of 50 is deliberately small. Items dropped by minPoints or minComments, duplicates removed inside a run, and diagnostic rows are not charged. A query that matches nothing writes one diagnostic row and costs you nothing.

๐Ÿ“ฅ What you give it

{
"query": "rust async",
"tags": "story",
"sortBy": "date",
"minPoints": 50,
"maxItems": 200
}
FieldDefaultWhat it is
querynoneKeywords. The form starts with openai in the box so it runs out of the box, but the API sends nothing unless you set it. Empty means the newest items for the type you picked.
tagsstorystory, show_hn, ask_hn, comment or front_page.
sortByrelevancerelevance for best match, date for newest first.
minPointsnoneOnly items with at least this many points. Applied before the row reaches you, so filtered items are not billed.
minCommentsnoneOnly items with at least this many comments, applied the same way. Refused with the comment type, because a comment has no comment count of its own.
maxItems501 to 1,000 per run. This is the dial that sets what a run costs.
notionConnectornoneOptional. Write every item into Notion as a page when the run finishes. Authorize the connector once under Settings, API and Integrations, MCP connectors, then pick it here.
notionParentIdnoneOptional. The Notion data source to write into. Empty creates the pages privately in your workspace.
proxyConfigurationoffOptional and not needed for a normal run.

๐Ÿ“ค What you get back

A real row from a recent run:

{
"ok": true,
"objectId": "38309611",
"type": "story",
"title": "OpenAI's board has fired Sam Altman",
"url": "https://openai.com/blog/openai-announces-leadership-transition",
"author": "davidbarker",
"points": 5710,
"numComments": 2530,
"createdAt": "2023-11-17T20:28:50Z",
"text": "",
"hnUrl": "https://news.ycombinator.com/item?id=38309611"
}
FieldHow to read it
objectIdThe item's Hacker News id, as a string. Stable, so use it as your dedupe key across runs.
typestory, show_hn, ask_hn, poll or comment.
title, urlThe item's own on a story. On a comment row both belong to the parent story. url is null on a text post that links nowhere.
textThe post or comment body with the HTML stripped out. Empty string on a plain link story, the post body on an Ask HN.
points, numCommentsOften null on comment rows, because comments carry no score in Hacker News search.
createdAtWhen the item was posted, in UTC.
hnUrlThe thread on news.ycombinator.com, built from objectId.

๐Ÿงพ Reading the output

Data rows carry ok: true. When something goes wrong, or a search matches nothing, you get one row with ok: false and an errorCode instead of an empty dataset, and that row is not charged.

CodeWhat it means
NO_RESULTSThe search matched nothing. Check the spelling, or widen it by clearing minPoints or minComments.
BAD_INPUTThe input asks for something the search cannot do, like a comment minimum on comments. Nothing was searched; the row's note says what to change.
RATE_LIMITEDToo many requests in a short window. Wait a little and run it again.
BLOCKEDThat search could not be run this time. Re-run it.
NOT_FOUNDThe address the run asked for came back missing.
SERVER_ERRORThe search answered with an error of its own.
NETWORKThe search was unreachable or answered badly. Re-run it.

One more thing about the Console's overview table: it shows title, author, points, comments and the two links, and it has no column for text. So a comments run looks odd there, every row showing its parent story's title with no comment body. Switch to all fields, or export the JSON, and the text is there.

๐Ÿ’ก What people use it for

  • Mention tracking. Put your product name in the query, set the type to comments, sort by date and run it every few hours, deduping on objectId. Threads move fast and hearing about yours from a customer is worse.
  • A feed of new launches: Show HN, sorted by date, no query needed.
  • The monthly "Who is hiring" lists. Those are comments under a post rather than the post itself, so search hiring with the type set to comments.
  • A front-page digest on a schedule, with the query box cleared.
  • Checking how a topic has been received over the years, sorted by points.

๐Ÿšง What it does not do

  • 1,000 rows per run. Split a bigger job across several queries or date ranges.
  • No full threads. You get the comments that match your search, not every reply under a story.
  • No user profiles, no karma totals, no submission history for a person.
  • The front page is filtered by your query too. For a plain digest, clear the query.
  • A minimum score does nothing useful on comments. Comment rows carry no score, so a run with both usually comes back with nothing at all. A minimum number of comments is refused on comments outright, for the same reason.
  • Scores are a snapshot. A story read an hour after it was posted has the points it had then.
  • No comment counts on comments. Those fields are for stories.
  • Relevance ordering is Hacker News search's own, and it is not something this actor can change.
  • It is not an official Hacker News tool. Dami's Studio is independent and is not affiliated with or endorsed by Hacker News or Y Combinator, or by any other company named on this page.

๐Ÿงญ Which developer scraper do you need?

If you wantUse
Articles from DEV.to by tag or author, with the full text on requestDEV.to Scraper
Hacker News stories, Ask and Show HN posts, and commentsThis one
Stack Overflow questions with the body and the accepted answerStack Overflow Scraper
Repositories and developer profiles from a GitHub searchGitHub Scraper
npm and PyPI package details and download countsnpm + PyPI Package Scraper
Product launches and their upvotesProduct Hunt Scraper
Posts from a Substack publicationSubstack Publication Scraper

โ“ Questions people ask

Do I need a Hacker News account or an API key?

Neither. Nothing to apply for.

What happens when my Hacker News search matches nothing?

You get one row with ok: false and errorCode: NO_RESULTS, and you are not charged for it.

Why did my front-page run come back nearly empty?

The query is applied to the front page as well, and the form starts with openai in the box. Clear it.

Can I pull a whole Hacker News comment thread?

No. Search returns matching comments, not the tree under a story. The hnUrl on each row opens the thread if you want to read it.

Can I call it from code or connect it to an AI assistant?

Yes. The API tab has ready-made code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/hacker-news-scraper. Either way the run happens on your Apify account at the same price.

These are public posts on a public site. Usernames are personal data under GDPR and similar laws when they point to a person, so have a reason for collecting them. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.

๐Ÿ†˜ If something breaks

Open the Issues tab on the actor page. Send the query you used and the run ID. The errorCode on the diagnostic row usually names the problem on its own.