# npm Scraper - Packages, Deps and Downloads (`s-r/npm-scraper`) Actor

Look up npm packages by name or search the registry. Returns version, licence, dependencies, release history, maintainers, deprecation status and real download counts from npm's own public APIs.

- **URL**: https://apify.com/s-r/npm-scraper.md
- **Developed by:** [SR](https://apify.com/s-r) (community)
- **Categories:** Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 run start fees

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## npm Scraper

Look up **npm packages** by name, or search the whole registry. You get the
version, licence, declared dependencies, full release history, maintainers,
deprecation status and **real download counts**.

No key, no login. This reads npm's own public APIs.

### The question this answers that a package page does not

**"Is this dependency actually alive?"**

`request` is the example worth keeping in mind. It has been deprecated for
years. It still returns a complete package document, a valid version, a licence
and a repository link — and **56 million downloads a month**. Nothing about the
response looks wrong.

npm never deletes a package, it flags it. So a scraper that reads the document
and reports what it finds will tell you a dead dependency is healthy. This Actor
lifts `deprecated` and `deprecated_reason` onto every row, and the run summary
counts how many of your packages are flagged.

The companion signal is `days_since_last_release`. Sort your dependency tree by
it and what has been abandoned under you becomes obvious.

### Three endpoints, because each holds something the others do not

| Endpoint | What only it has |
|---|---|
| `registry.npmjs.org/<pkg>` | Every version, dependencies, licence, maintainers, release dates |
| `registry.npmjs.org/-/v1/search` | The composite **score** and the **dependents** count |
| `api.npmjs.org/downloads` | Actual download numbers |

So the two modes give you different columns, and the run summary says which you
are in rather than leaving you to wonder why a field is empty.

**By name** returns the full document:

```
packages: ["express", "lodash", "request"]
```

**By search** returns scores and dependents:

```
search: "http client"
```

Download counts are added in either mode, from the third endpoint, one request
per package.

### A note on npm's scores

npm publishes a `score` object with a composite `final` value and a `detail`
breakdown of quality, popularity and maintenance.

**The breakdown is dead.** Verified across eight search results where the
composite ranged from 234 to 2,445: quality, popularity and maintenance came
back as exactly `1` for every single package. Three columns that are always 1
look like signal and are not, so this Actor does not return them.

What it does return is `score`, which genuinely varies, and `dependents`, which
is the count of packages depending on this one — 215,361 for `react`, 44 for a
niche HTTP client. That number is the honest popularity measure.

### Fields

- **Identity**: `name` (scoped names included), `version`, `description`,
  `keywords`, `url`
- **Legal**: `license`
- **Links**: `homepage`, `repository` (normalised from `git+https://…​.git` to a
  plain https URL)
- **Dependencies**: `dependencies`, `dependency_count`,
  `dev_dependency_count`, `peer_dependency_count`, `engines`
- **People**: `maintainers`, `publisher`
- **History**: `version_count`, `first_release_at`, `last_release_at`,
  `days_since_last_release`
- **Health**: `deprecated`, `deprecated_reason`
- **Artefact**: `tarball`, `unpacked_size_bytes`, `dist_tags`
- **Popularity**: `downloads_month` (or day/week), `score`, `dependents`

### Input reference

| Field | Type | Default |
|---|---|---|
| `packages` | list of exact names | `["express","lodash","request"]` |
| `search` | registry search term | — |
| `include_downloads` | fetch real download counts | `true` |
| `downloads_period` | last-day, last-week, last-month | `last-month` |
| `limit` | 1-2000 | 50 |
| `retries` | 1-6 | 3 |

### A note on speed

Package documents are large. `express` is 805 KB because it carries all 288
versions; `@types/node` has 2,359 versions. That weight is the cost of a
by-name lookup, so concurrency is capped at five to stay polite to a registry
that serves the whole ecosystem for free.

If you only need an overview rather than dependency detail, the search mode is
far lighter and gives you scores and dependents on top.

### Typical uses

- **Dependency audit.** Feed your `package.json` dependencies in and get every
  licence, deprecation flag and last-release date back in one table.
- **Supply-chain review.** `dependency_count` and the `dependencies` list show
  how much a package drags in. `maintainers` shows how many people can publish.
- **Abandonment check.** Sort by `days_since_last_release`, filter on
  `deprecated`.
- **Ecosystem research.** Search a term and rank by `dependents` to see what the
  ecosystem actually builds on, rather than what markets itself best.
- **Package selection.** Compare candidates on downloads, dependents, dependency
  weight and release recency in one run.

### Notes

A name that does not exist returns a `not_found` error naming it, never a row of
nulls. A 404 is definitive and is not retried; a 500 or timeout is retried with
backoff and only then reported. You always know which you are looking at.

`license` is returned only when npm publishes it as a plain string. Some
packages declare it as an object or an SPDX expression array, and those come
back `null` rather than being flattened into something that looks canonical and
is not.

### Scoped packages

Scoped names (`@types/node`, `@actions/http-client`) work everywhere a plain
name does. The slash is preserved through the request and the resulting `url`
points at the right page, which is worth mentioning because it is the detail
that most quickly-written npm scrapers get wrong: a naive URL-encode turns
`@types/node` into `%40types%2Fnode` and the registry answers 404.

### What this Actor does not do

**No tarball download or contents.** `tarball` gives you the URL and
`unpacked_size_bytes` the size, but the package contents are not fetched. That
is a different job and a much heavier one.

**No vulnerability data.** npm's audit endpoint needs a different request shape
and returns advisories rather than package metadata. Deprecation is published in
the registry and is returned; CVEs are not.

**No dependency tree.** You get a package's own declared dependencies, not the
resolved tree beneath them. Feeding the dependency names back in as `packages`
gives you the next level, which is usually the practical way to walk a tree a
layer at a time.

# Actor input Schema

## `packages` (type: `array`):

Exact npm package names to look up, scoped names included. This path returns the full package document: dependencies, release history and deprecation.

## `search` (type: `string`):

Search the registry instead of naming packages. This path also returns npm's search score and the dependents count, which the package document does not carry.

## `include_downloads` (type: `boolean`):

Fetch real download counts from npm's downloads API. One extra request per package.

## `downloads_period` (type: `string`):

Which period the download count covers.

## `limit` (type: `integer`):

How many packages to return.

## `retries` (type: `integer`):

Retries with backoff before a request is reported as an error.

## Actor input object example

```json
{
  "packages": [
    "express",
    "lodash",
    "request"
  ],
  "search": "http client",
  "include_downloads": true,
  "downloads_period": "last-month",
  "limit": 50,
  "retries": 3
}
```

# Actor output Schema

## `packages` (type: `string`):

One row per npm package.

## `summary` (type: `string`):

Counts, deprecated packages and download coverage.

## `errors` (type: `string`):

Failures with a code and a redacted message.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "packages": [
        "express",
        "lodash",
        "request"
    ],
    "limit": 50,
    "retries": 3
};

// Run the Actor and wait for it to finish
const run = await client.actor("s-r/npm-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "packages": [
        "express",
        "lodash",
        "request",
    ],
    "limit": 50,
    "retries": 3,
}

# Run the Actor and wait for it to finish
run = client.actor("s-r/npm-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "packages": [
    "express",
    "lodash",
    "request"
  ],
  "limit": 50,
  "retries": 3
}' |
apify call s-r/npm-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,s-r/npm-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/cSEYJIEQwMxDOf2DI/builds/GSL8Ml6RKXpTdfvcW/openapi.json
