Hacker News Tech Stories Scraper avatar

Hacker News Tech Stories Scraper

Pricing

from $2.99 / 1,000 technology stories

Go to Apify Store
Hacker News Tech Stories Scraper

Hacker News Tech Stories Scraper

Extracts public technology-focused Hacker News stories from Algolia with rich ranking, highlights and optional comments.

Pricing

from $2.99 / 1,000 technology stories

Rating

0.0

(0)

Developer

w3crawler

w3crawler

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Searches public Hacker News stories through the Algolia HN search API, using a configurable technology query such as technology, AI, programming, or open source. Records include stable story IDs, titles, authors, normalized timestamps, points, comment counts, canonical HTTPS Hacker News URLs, optional HTTPS story URLs, domains, story text, query/rank metadata, tags, Algolia highlights, child IDs, and optional real top-level Firebase comments.

Input

{
"query": "technology",
"maxItems": 25,
"sortByDate": false,
"includeComments": true,
"maxCommentsPerStory": 2
}

Requests are bounded, retried and delayed; deleted or incomplete comments are omitted, insecure story URLs are not followed or emitted, and no search result is fabricated. Search failures are represented by a fail-closed diagnostic row instead of an unstructured process error. Optional comment failures remain attached to the affected story and mark the run partial.

Output

Each dataset item contains a stable recordId and recordType (story or diagnostic). Story items contain Algolia story identity, title, author, timestamps, points, comment count, a canonical Hacker News URL, an optional HTTPS story URL, domain, text, query and rank metadata, tags, highlights, child IDs, and optional Firebase topComments. Diagnostic items contain found: false, dataAvailable: false, and structured access/availability details.

Example input: {"query":"technology","maxItems":10,"sortByDate":false,"includeComments":true,"maxCommentsPerStory":2}

Example output: {"storyId":"12345678","titleText":"Example technology story","author":"example-user","points":180,"commentCount":32,"storyUrl":"https://example.com/story","searchQuery":"technology","searchRank":1,"tags":["story","author_example-user"],"hackerNewsUrl":"https://news.ycombinator.com/item?id=12345678"}

Key-value store

The exact OUTPUT key contains the JSON run summary, including status, requested and stored counts, duplicate/skipped counts, public source URL, and comment-enrichment metrics.

Cost and limits

The main cost is one Algolia request plus up to four bounded Firebase comment requests per story. Keep maxItems and maxCommentsPerStory modest.

Local run

apify run --purge --entrypoint src/main.js --input '{"query":"technology","maxItems":5,"includeComments":true,"maxCommentsPerStory":2}'
npm run validate

The validator checks stable IDs, typed records, HTTPS evidence URLs, timestamps, null-free output, duplicate story identities, comment diagnostics, and consistency between the dataset and OUTPUT.

QA validated locally and on Apify Cloud with a bounded live technology-story search. Generated by Codex with GPT-5.6 Sol.