# Product Hunt Launches Scraper (`scrapyx/producthunt-launches-scraper`) Actor

Ranked Product Hunt launches by day, week, month or year, plus product detail. Flags the paid placements injected into the ranking, and refuses impossible dates that Product Hunt would silently serve as a different day.

- **URL**: https://apify.com/scrapyx/producthunt-launches-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** Business, Marketing, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.40 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Product Hunt Launches Scraper

Ranked **Product Hunt** launches by day, week, month or year, plus full
product detail. HTTP-only, no API key, no OAuth, no login, no browser.

### Modes

| Mode | What you get |
| --- | --- |
| `leaderboard` | Ranked launches for a period. `daily` walks a date range one day at a time; `weekly`, `monthly` and `yearly` fetch one board. Each row carries rank, votes, comments, tagline and topics. |
| `products` | The full product document for slugs you name — description, category, makers, screenshots, overall rating and the pros/cons lists. |

**There is no keyword-search mode, by design.** Product Hunt's `robots.txt`
disallows `/search*` outright, so this actor does not offer it. Same for user
profiles (`/@*/*`).

### Six upstream quirks it corrects

#### 1. `robots.txt` itself answered 403 — and that was not the policy

The first fetch of `producthunt.com/robots.txt` came back **HTTP 403**. Running
the TLS ladder over the robots file itself showed why:

| Profile | robots.txt |
| --- | --- |
| chrome131, safari17\_0, safari18\_0, safari17\_2\_ios, firefox133 | **200** |
| chrome124, chrome136, edge101, chrome99\_android | **403** (Cloudflare) |

The 403 was a Cloudflare challenge, not a policy statement. The real file
**names no AI crawler at all**. A 403 on robots.txt is a transport problem
until proven otherwise — reading it as a verdict would have discarded this
target entirely.

#### 2. An impossible date is silently served as a **different real day**, and both echoes lie

```
/leaderboard/daily/2026/13/99  -> HTTP 200, heading "April 9, 2027"
/leaderboard/daily/2026/2/31   -> HTTP 200, heading "March 3, 2026"
```

February 31st does not exist. Product Hunt rolls it forward rather than
refusing. Neither available signal catches it:

- **`<link rel="canonical">` echoes your bogus input back verbatim**
  (`/daily/2026/13/99`), so the canonical-gating trick that works on Land.com,
  Immoweb and Otodom does nothing here.
- **The heading lies too.** `2026/2/31` claims "March 3, 2026" while sharing
  only **1 of 19** products with the genuine March 3 page.

With no trustworthy server signal, the only correct defence is to **refuse
impossible dates before spending a request**. This actor validates every date
locally and fails loudly. Out-of-range `month` and `week` values are refused
the same way.

#### 3. Paid placements are injected into the ranking, and the obvious marker is useless

One daily board's vote counts, in list order:

```
388 302 235 214 182 160 145 140 132 [911] 120 115 113 112 109 [696] 106 96 93
```

The two bracketed entries break the descending order. They are ad slots.
Grepping the card markup for the word **"Promoted" matches all 19 cards** — it
lives in shared markup, so it is a false positive on every one.

The only honest discriminator is the **numeric rank prefix**: ranked entries
read `"1. HyNote for Mac"`, promoted ones read `"Framer AI Agents"` with no
prefix. Promoted rows are **kept and flagged** (`isPromoted`, `rank: null`) —
which of the two you want is your decision. Set `includePromoted: false` to
drop them; they stay counted in the summary either way.

#### 4. …except the **yearly** board numbers nothing at all

| Board | cards | numbered |
| --- | ---: | ---: |
| daily | 19 | 17 |
| weekly | 19 | 17 |
| monthly | 19 | 17 |
| **yearly** | 19 | **0** |

The yearly board prints no rank numbers, and its votes descend cleanly
(1,839 / 1,673 / 1,571 / 1,417 …) — all 19 are genuine. A flat "no prefix means
promoted" rule labelled **every one of them an ad**.

So promoted detection is decided **per page**: if a board numbers some of its
cards, the unnumbered ones are ads; if it numbers none, the prefix carries no
signal and detection is reported as **unavailable**
(`promotedDetectionAvailable: false`, `pagesWithoutPromotedDetection`) rather
than guessed. On such a board `rank` falls back to list position and says so
via `rankFromListPosition: true`.

#### 5. Rank does **not** follow vote count

| Board | inversions |
| --- | --- |
| daily | 1 — rank 16 has 96 votes, above rank 15's 92 |
| monthly | 5, starting at the top — **rank 1 has 838 votes, rank 2 has 945** |

Product Hunt's ordering is not a vote sort. Re-deriving rank from `votes`
produces a different order than the site shows, so every row carries
`rankFollowsVotes: false` and the actor never implies otherwise.

#### 6. `positiveNotes` and `negativeNotes` can be byte-identical

On one product both lists were the **same 22 entries** — "all-in-one
workspace", "team collaboration", "project management" … presented as both the
pros and the cons. Publishing them as pros/cons would ship a confidently wrong
cons column. Both are emitted verbatim plus `notesListsAreIdentical`, and the
summary counts how many products collided, so you can see it rather than
inherit it.

### Output

One `SEARCH_SUMMARY` per run, one `LAUNCH` per leaderboard entry, one `PRODUCT`
per product document, one `ERROR` per failure.

`LAUNCH`: `productSlug`, `productUrl`, `productName`, `tagline`, `rank`,
`listPosition`, `rankFromListPosition`, `isPromoted`,
`promotedDetectionAvailable`, `rankFollowsVotes`, `votes`, `comments`,
`topics`, `period`, `periodRequested`, `pageHeading`.

`PRODUCT`: `productName`, `description`, `applicationCategory`,
`operatingSystem`, `datePublished`, `dateModified`, `imageUrl`, `screenshots`,
`ratingValue`, `ratingCount`, `makers`, `positiveNotes`, `negativeNotes`,
`notesListsAreIdentical`.

### Limits

- **No search and no user profiles** — `robots.txt` disallows both.
- **No topic mode.** `/topics/{slug}` is policy-clean and returns 200, but it
  renders zero leaderboard cards and its product links are polluted by
  navigation and footer entries (one probe picked up `attio/reviews?ref=footer`
  as if it were a product). Declared out of scope rather than shipped
  half-reliable.
- **Slugs cannot be guessed** from a product's name — a wrong guess is an
  honest 404. Take them from this actor's own leaderboard output.
- A board holds roughly 17 ranked entries plus any promoted slots; there is no
  deeper pagination on a leaderboard.

# Actor input Schema

## `mode` (type: `string`):

leaderboard = ranked launches for a day, week, month or year. products = full detail for product slugs you name. There is no keyword-search mode: Product Hunt's robots.txt disallows /search\* outright.

## `period` (type: `string`):

daily walks a date range one day at a time. weekly/monthly/yearly fetch one board.

## `startDate` (type: `string`):

YYYY-MM-DD. Impossible dates are refused here rather than sent: Product Hunt answers HTTP 200 for something like 2026-02-31 and silently serves a DIFFERENT real date, and neither its canonical URL nor its page heading reveals the swap.

## `endDate` (type: `string`):

YYYY-MM-DD. Defaults to the start date. Max 400 days per run.

## `year` (type: `integer`):

Product Hunt's data starts in 2013.

## `month` (type: `integer`):

1-12. Out-of-range values are refused locally, because upstream rolls them into a different real period without saying so.

## `week` (type: `integer`):

1-53.

## `productSlugs` (type: `array`):

For mode='products'. A slug like `notion` or a full producthunt.com/products/ URL. Slugs cannot be guessed from a product's name — a wrong guess is an honest 404 — so take them from this actor's own leaderboard output.

## `includePromoted` (type: `boolean`):

Product Hunt injects paid placements into the ranking. They are always FLAGGED (rank is null, isPromoted is true) — this only controls whether they are emitted at all. On by default so nothing is hidden from you.

## `fetchProductDetails` (type: `boolean`):

Leaderboard mode only — products mode always fetches it. Adds description, category, maker list, screenshots and the overall rating, at one extra request per unique product.

## `maxResults` (type: `integer`):

Set 0 for unlimited. A daily board holds roughly 17 ranked entries plus any promoted slots.

## `maxConcurrency` (type: `integer`):

Product-detail fetches in flight at once. Leaderboard days are sequential.

## `minRequestInterval` (type: `integer`):

Politeness pacing shared across all workers. 0 uses the built-in default of 1s. Raising this is the right response to Cloudflare challenges.

## `proxyConfiguration` (type: `object`):

Cloudflare here gates the TLS fingerprint rather than the IP — the whole reconnaissance ran proxy-free. Residential is still the default for cloud runs, because datacenter ASNs are scored separately.

## Actor input object example

```json
{
  "mode": "leaderboard",
  "period": "daily",
  "startDate": "2026-08-25",
  "endDate": "2026-08-25",
  "year": 2026,
  "month": 7,
  "week": 33,
  "productSlugs": [
    "notion",
    "https://www.producthunt.com/products/raycast"
  ],
  "includePromoted": true,
  "fetchProductDetails": false,
  "maxResults": 200,
  "maxConcurrency": 3,
  "minRequestInterval": 0,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/producthunt-launches-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/producthunt-launches-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call scrapyx/producthunt-launches-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/producthunt-launches-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/SVICrLa3tGu32xdcV/builds/2BhuPpagUBOtl1rwr/openapi.json
