# npm & PyPI Package Stats Scraper (`67-labs/npm-pypi-package-stats-scraper`) Actor

Get npm and PyPI package stats: latest version, license, weekly and monthly downloads, maintainers, last publish, repo. CSV or Sheets.

- **URL**: https://apify.com/67-labs/npm-pypi-package-stats-scraper.md
- **Developed by:** [Mokksh Bhatt](https://apify.com/67-labs) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 result (one package)s

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## npm & PyPI Package Stats Scraper — downloads, versions, license, maintainers and last publish for any package

Get the public facts of any npm or PyPI package in one row: latest version, license, weekly and monthly downloads, number of maintainers, number of versions, when it was created and last published, dependencies, the repository link, and whether it is deprecated. Give it package names or package links. Export to CSV, JSON, Excel or Google Sheets. It uses the registries' own open APIs, so the numbers are official and the actor is stable.

**Who uses this:** developers and engineering leads choosing libraries, security and open source teams auditing dependencies, developer-tools companies tracking adoption, investors and analysts measuring open source projects, developer relations teams comparing themselves with competitors, AI agents that need current package facts.

### What you get per package

| Field | What it is |
|---|---|
| `ecosystem`, `name`, `url` | `npm` or `pypi`, the package name, and its page |
| `description`, `keywords`, `author` | What the package says about itself |
| `latestVersion`, `versionsCount` | Newest version, and how many versions were ever published |
| `license` | License name or SPDX expression |
| `homepage`, `repositoryUrl` | Website, and the source repository as a clean `https://` link |
| `weeklyDownloads`, `monthlyDownloads` | Downloads in the last 7 and 30 days |
| `dailyDownloads` | Downloads yesterday (PyPI only) |
| `maintainers`, `maintainersCount` | Registry accounts that can publish (npm: maintainers, PyPI: owners) |
| `createdAt`, `lastPublishedAt` | First publish, and the latest release |
| `dependencyCount` | Direct dependencies of the latest version |
| `requires` | Required runtime, for example `node >= 18` or `python >=3.10` |
| `deprecated`, `deprecatedMessage` | `true` if the latest version is deprecated (npm) or yanked (PyPI), and the message |
| `vulnerabilitiesCount` | Known vulnerabilities of the latest version (PyPI only) |
| `scrapedAt` | When the row was collected |

### Price

**$0.01 per run + $1 per 1,000 packages.** That is two pay-per-event charges: `actor-start` ($0.01) once per run and `result` ($0.001) per package. No compute fees on top. The free Apify plan ($5 per month credit) covers about 4,900 packages a month at no cost. A package that does not exist is not charged. A run where the registries do not answer is not charged.

### How to use

1. Open the Input tab and add packages, one per line. A plain name such as `react` is looked up on npm. Write `pypi:requests` for PyPI, or paste the package link. Scoped npm names like `@types/node` work.
2. Choose **Registry for plain names** if you want plain names looked up on PyPI or on both.
3. Click Start. When it is done, open the Output tab and export the table as CSV, JSON or Excel, or send it to Google Sheets with Apify's Google Sheets integration.

### Input examples

Three npm packages and two Python packages:

```json
{ "packages": ["react", "lodash", "@types/node", "pypi:requests", "pypi:numpy"] }
```

Package links work too:

```json
{ "packages": ["https://www.npmjs.com/package/express", "https://pypi.org/project/flask/"] }
```

The same names on both registries, without download numbers, for a fast check of who owns a name:

```json
{ "packages": ["pandas", "requests"], "ecosystem": "both", "includeDownloads": false }
```

### Output example

```json
{
    "ecosystem": "npm",
    "name": "express",
    "url": "https://www.npmjs.com/package/express",
    "description": "Fast, unopinionated, minimalist web framework",
    "latestVersion": "5.2.1",
    "license": "MIT",
    "homepage": "https://expressjs.com/",
    "repositoryUrl": "https://github.com/expressjs/express",
    "keywords": ["express", "framework", "sinatra", "web"],
    "author": "TJ Holowaychuk",
    "maintainers": ["wesleytodd", "jonchurch", "ctcpip"],
    "maintainersCount": 5,
    "versionsCount": 289,
    "createdAt": "2010-12-29T19:38:25.450Z",
    "lastPublishedAt": "2025-12-01T20:49:43.268Z",
    "weeklyDownloads": 130322863,
    "monthlyDownloads": 457752996,
    "dailyDownloads": null,
    "deprecated": false,
    "deprecatedMessage": null,
    "dependencyCount": 28,
    "requires": "node >= 18",
    "vulnerabilitiesCount": null,
    "scrapedAt": "2026-09-26T10:00:00.000Z"
}
```

The lists of keywords and maintainers are shortened in this example.

### Limits

- Download numbers come from `api.npmjs.org` (npm) and `pypistats.org` (PyPI). They are public statistics that count downloads, including automated ones such as build systems, not unique people.
- For PyPI, `maintainers` are the owners the registry lists. Packages owned by an organization can show 0.
- `vulnerabilitiesCount` is available for PyPI only. For npm it is empty.
- `pypistats.org` limits how fast it can be asked. The actor waits and tries again. If a number is still missing after that, the field is empty and the row is still saved.
- A package that does not exist is named in the run log and skipped.

### FAQ

**Can I track adoption over time?** Yes. Schedule the actor in Apify with the same packages and compare `weeklyDownloads` between runs.

**Can I compare a package with its competitors?** Yes. Put all of them in one run and sort the table by `monthlyDownloads`.

**Can an AI agent use it?** Yes. It works through Apify's API and MCP server. Give it a list of packages.

**Something is wrong. What now?** Open an issue on the Issues tab with the package and the input you used. Field requests are welcome.

### Changelog

- **0.1** First release: npm and PyPI stats, downloads, maintainers, versions, repository, pay per event.

# Actor input Schema

## `packages` (type: `array`):

One entry per package: a name (react), a name with the registry (npm:express, pypi:requests), or the package link (https://www.npmjs.com/package/react, https://pypi.org/project/requests/). Scoped names work: @types/node.

## `ecosystem` (type: `string`):

Where to look for names that have no npm: or pypi: in front. Both looks in both registries and gives a row for each one found.

## `includeDownloads` (type: `boolean`):

Adds weekly and monthly downloads (and daily for PyPI). Needs 1 or 2 more requests per package.

## `proxyConfiguration` (type: `object`):

Leave off. The registries are open.

## Actor input object example

```json
{
  "packages": [
    "react",
    "npm:@types/node",
    "pypi:numpy",
    "https://pypi.org/project/requests/"
  ],
  "ecosystem": "npm",
  "includeDownloads": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `packages` (type: `string`):

One row per package: version, license, downloads, maintainers, versions, last publish, repository.

## `packagesCsv` (type: `string`):

Same rows as CSV for Google Sheets or Excel.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "packages": [
        "react",
        "pypi:requests"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("67-labs/npm-pypi-package-stats-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "packages": [
        "react",
        "pypi:requests",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("67-labs/npm-pypi-package-stats-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "packages": [
    "react",
    "pypi:requests"
  ]
}' |
apify call 67-labs/npm-pypi-package-stats-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,67-labs/npm-pypi-package-stats-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/3oRJMsBWT6vmFdjp5/builds/QMIfqQBB61blbOLlk/openapi.json
