# Package Registry Scraper — npm, PyPI & crates.io (`hichemdev/package-registry-scraper`) Actor

Scrape software packages from npm, PyPI, crates.io, RubyGems and Packagist through one unified schema: version, description, licence, downloads, dependents, dependency and version counts, maintainers, repository, publish dates and deprecation status. No API key.

- **URL**: https://apify.com/hichemdev/package-registry-scraper.md
- **Developed by:** [Hichem Ben Moussa](https://apify.com/hichemdev) (community)
- **Categories:** Developer tools, Business, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 packages

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Package Registry Scraper — npm, PyPI, crates.io & More

Scrape software packages from **five registries through one unified schema**: npm (JavaScript), PyPI (Python), crates.io (Rust), RubyGems (Ruby) and Packagist (PHP).

Downloads, licences, dependency and version counts, maintainers, repository links, publish dates and deprecation status — the same field names whatever the ecosystem, so one pipeline handles all five.

No API key required.

### What you get

| Field | Description |
|---|---|
| `registry` | Which registry the row came from |
| `name`, `version` | Package identity and current version |
| `description` | Summary |
| `license` | Licence identifier |
| `homepage`, `repository` | Links, with git URLs normalised to browsable HTTPS |
| `author`, `maintainers`, `maintainerCount` | **A single maintainer on a widely used package is a supply-chain risk signal** |
| `downloadsLastWeek`, `downloadsLastMonth`, `downloadsTotal` | Whichever the registry publishes |
| `dependents` | How many packages depend on this one |
| `dependencyCount` | How many it pulls in itself |
| `versionCount` | Release count |
| `publishedAt`, `lastPublishedAt` | First and latest release — the staleness signal |
| `stars`, `forks`, `openIssues` | Repository stats, where the registry exposes them |
| `requiresRuntime` | Node/Python/PHP/Rust version constraint |
| `isDeprecated`, `deprecationReason` | Deprecated or yanked, and why |
| `url` | Package page |

### Coverage by registry

Registries publish different things, and this actor reports what each one actually has rather than inventing the rest:

| | npm | PyPI | crates.io | RubyGems | Packagist |
|---|---|---|---|---|---|
| Keyword search | ✅ | ❌ | ✅ | ✅ | ✅ |
| Weekly / monthly downloads | ✅ | ❌ | monthly | ❌ | monthly |
| Lifetime downloads | ❌ | ❌ | ✅ | ✅ | ✅ |
| Dependents count | ✅ | ❌ | ❌ | ❌ | ✅ |
| Repository stars | ❌ | ❌ | ❌ | ❌ | ✅ |
| Version count | ❌ | ✅ | ✅ | ❌ | ✅ |

**PyPI has retired its public search API** — it now returns an HTML page, not JSON. So for PyPI you must name the packages you want in *Specific packages*; a search term there returns nothing and the actor says so in the log instead of failing quietly. PyPI also publishes no download counts through its JSON API.

### Example input

```json
{
  "registry": "npm",
  "searchTerm": "http client",
  "minDownloads": 100000,
  "maxPackages": 100
}
```

Audit an exact dependency list, on any registry:

```json
{ "registry": "pypi", "packages": ["requests", "urllib3", "certifi"] }
{ "registry": "npm", "packages": ["express", "@types/node", "lodash"] }
{ "registry": "crates", "packages": ["serde", "tokio"] }
```

### Example output

```json
{
  "registry": "npm",
  "name": "express",
  "version": "5.2.1",
  "description": "Fast, unopinionated, minimalist web framework",
  "license": "MIT",
  "repository": "https://github.com/expressjs/express",
  "maintainers": ["wesleytodd", "jonchurch"],
  "maintainerCount": 2,
  "downloadsLastWeek": 158867531,
  "downloadsLastMonth": 536200734,
  "dependencyCount": 31,
  "isDeprecated": false,
  "url": "https://www.npmjs.com/package/express"
}
```

### Who uses this

- **Developer-tool marketing** — size your category: which packages in "http client" have real adoption, and which are abandoned
- **Supply-chain security** — flag dependencies that are deprecated, single-maintainer, or have not shipped in years
- **Licence compliance** — pull the licence for every package in your manifest across all five ecosystems in one run
- **Open-source maintainers** — track your download curve against direct competitors weekly
- **Technical due diligence** — assess a target's dependency health before an acquisition
- **Package comparison sites** — populate a directory with real numbers rather than estimates

The staleness screen is the most valuable one: `lastPublishedAt` more than a year old plus a high `dependents` count is exactly the shape of the dependencies that cause incidents.

### Notes

- **npm's per-package document is ~800 KB** because it embeds every version and the readme. The actor reads the compact `/latest` view and fetches download counts separately, which keeps runs fast.
- **Download windows are not comparable across registries.** npm reports a rolling week and month; crates.io and Packagist report monthly plus lifetime; RubyGems reports lifetime only. Compare within a registry, not across them, and use `minDownloads` knowing it tests the best figure available for that row.
- **Search results are ranked by the registry's own relevance or download sort**, not re-ranked here.
- `isDeprecated` covers npm's `deprecated` field, PyPI and crates.io yanks, and RubyGems yanks — different mechanisms, one boolean.
- This is an unofficial actor and is not affiliated with npm, the PSF, the Rust Foundation, RubyGems or Packagist.

### Pricing

Pay per result. Each package returned counts as one result, and the download filter is applied before charging.

# Actor input Schema

## `registry` (type: `string`):

Which package registry to read.

## `packages` (type: `array`):

Exact package names, e.g. express, @types/node, requests, serde. Required for PyPI, which no longer offers a search API.

## `searchTerm` (type: `string`):

Search the registry instead of naming packages. Supported on npm, crates.io, RubyGems and Packagist.

## `minDownloads` (type: `integer`):

Keep only packages whose best available download figure is at least this high.

## `maxPackages` (type: `integer`):

Stop after this many packages.

## `proxyConfiguration` (type: `object`):

Optional. A proxy is not required for this actor.

## Actor input object example

```json
{
  "registry": "npm",
  "searchTerm": "http client",
  "maxPackages": 50,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

All scraped items as JSON.

## `overview` (type: `string`):

Browse results in the Apify Console.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "registry": "npm",
    "searchTerm": "http client",
    "maxPackages": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("hichemdev/package-registry-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "registry": "npm",
    "searchTerm": "http client",
    "maxPackages": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("hichemdev/package-registry-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "registry": "npm",
  "searchTerm": "http client",
  "maxPackages": 50
}' |
apify call hichemdev/package-registry-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,hichemdev/package-registry-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/BXMVZMCWTBRoBqaRN/builds/sjHQxZyaPHk5Bzduk/openapi.json
