npm Scraper - Packages, Deps and Downloads avatar

npm Scraper - Packages, Deps and Downloads

Pricing

from $1.00 / 1,000 run start fees

Go to Apify Store
npm Scraper - Packages, Deps and Downloads

npm Scraper - Packages, Deps and Downloads

Look up npm packages by name or search the registry. Returns version, licence, dependencies, release history, maintainers, deprecation status and real download counts from npm's own public APIs.

Pricing

from $1.00 / 1,000 run start fees

Rating

0.0

(0)

Developer

SR

SR

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Categories

Share

npm Scraper

Look up npm packages by name, or search the whole registry. You get the version, licence, declared dependencies, full release history, maintainers, deprecation status and real download counts.

No key, no login. This reads npm's own public APIs.

The question this answers that a package page does not

"Is this dependency actually alive?"

request is the example worth keeping in mind. It has been deprecated for years. It still returns a complete package document, a valid version, a licence and a repository link — and 56 million downloads a month. Nothing about the response looks wrong.

npm never deletes a package, it flags it. So a scraper that reads the document and reports what it finds will tell you a dead dependency is healthy. This Actor lifts deprecated and deprecated_reason onto every row, and the run summary counts how many of your packages are flagged.

The companion signal is days_since_last_release. Sort your dependency tree by it and what has been abandoned under you becomes obvious.

Three endpoints, because each holds something the others do not

EndpointWhat only it has
registry.npmjs.org/<pkg>Every version, dependencies, licence, maintainers, release dates
registry.npmjs.org/-/v1/searchThe composite score and the dependents count
api.npmjs.org/downloadsActual download numbers

So the two modes give you different columns, and the run summary says which you are in rather than leaving you to wonder why a field is empty.

By name returns the full document:

packages: ["express", "lodash", "request"]

By search returns scores and dependents:

search: "http client"

Download counts are added in either mode, from the third endpoint, one request per package.

A note on npm's scores

npm publishes a score object with a composite final value and a detail breakdown of quality, popularity and maintenance.

The breakdown is dead. Verified across eight search results where the composite ranged from 234 to 2,445: quality, popularity and maintenance came back as exactly 1 for every single package. Three columns that are always 1 look like signal and are not, so this Actor does not return them.

What it does return is score, which genuinely varies, and dependents, which is the count of packages depending on this one — 215,361 for react, 44 for a niche HTTP client. That number is the honest popularity measure.

Fields

  • Identity: name (scoped names included), version, description, keywords, url
  • Legal: license
  • Links: homepage, repository (normalised from git+https://…​.git to a plain https URL)
  • Dependencies: dependencies, dependency_count, dev_dependency_count, peer_dependency_count, engines
  • People: maintainers, publisher
  • History: version_count, first_release_at, last_release_at, days_since_last_release
  • Health: deprecated, deprecated_reason
  • Artefact: tarball, unpacked_size_bytes, dist_tags
  • Popularity: downloads_month (or day/week), score, dependents

Input reference

FieldTypeDefault
packageslist of exact names["express","lodash","request"]
searchregistry search term—
include_downloadsfetch real download countstrue
downloads_periodlast-day, last-week, last-monthlast-month
limit1-200050
retries1-63

A note on speed

Package documents are large. express is 805 KB because it carries all 288 versions; @types/node has 2,359 versions. That weight is the cost of a by-name lookup, so concurrency is capped at five to stay polite to a registry that serves the whole ecosystem for free.

If you only need an overview rather than dependency detail, the search mode is far lighter and gives you scores and dependents on top.

Typical uses

  • Dependency audit. Feed your package.json dependencies in and get every licence, deprecation flag and last-release date back in one table.
  • Supply-chain review. dependency_count and the dependencies list show how much a package drags in. maintainers shows how many people can publish.
  • Abandonment check. Sort by days_since_last_release, filter on deprecated.
  • Ecosystem research. Search a term and rank by dependents to see what the ecosystem actually builds on, rather than what markets itself best.
  • Package selection. Compare candidates on downloads, dependents, dependency weight and release recency in one run.

Notes

A name that does not exist returns a not_found error naming it, never a row of nulls. A 404 is definitive and is not retried; a 500 or timeout is retried with backoff and only then reported. You always know which you are looking at.

license is returned only when npm publishes it as a plain string. Some packages declare it as an object or an SPDX expression array, and those come back null rather than being flattened into something that looks canonical and is not.

Scoped packages

Scoped names (@types/node, @actions/http-client) work everywhere a plain name does. The slash is preserved through the request and the resulting url points at the right page, which is worth mentioning because it is the detail that most quickly-written npm scrapers get wrong: a naive URL-encode turns @types/node into %40types%2Fnode and the registry answers 404.

What this Actor does not do

No tarball download or contents. tarball gives you the URL and unpacked_size_bytes the size, but the package contents are not fetched. That is a different job and a much heavier one.

No vulnerability data. npm's audit endpoint needs a different request shape and returns advisories rather than package metadata. Deprecation is published in the registry and is returned; CVEs are not.

No dependency tree. You get a package's own declared dependencies, not the resolved tree beneath them. Feeding the dependency names back in as packages gives you the next level, which is usually the practical way to walk a tree a layer at a time.