npm + PyPI Package Scraper - Metadata & Downloads avatar

npm + PyPI Package Scraper - Metadata & Downloads

Pricing

from $1.94 / 1,000 packages

Go to Apify Store
npm + PyPI Package Scraper - Metadata & Downloads

npm + PyPI Package Scraper - Metadata & Downloads

Look up npm and PyPI packages by exact name, or search npm by keyword. Each row has the latest version, description, author, repository, licence, keywords and project link, plus monthly downloads on npm. No API key. PyPI matches exact names only. $2 per 1,000 packages.

Pricing

from $1.94 / 1,000 packages

Rating

0.0

(0)

Developer

Dami's Studio

Dami's Studio

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

1

Monthly active users

21 hours ago

Last modified

Share

Look up npm and PyPI packages by exact name, or search npm by keyword, and get one row per package: the latest version, description, author, homepage, repository, licence, keywords and the project page, plus last month's download count on npm. Both registries come back in the same row shape.

The two registries are not symmetrical, and it matters before you plan a run. npm does keyword search and exact names. PyPI does exact names only, because it publishes no clean search API for anyone to call. And an npm search row never carries the licence, so a licence audit goes through names.

InputExact package names on npm or PyPI, or keywords for an npm search
OutputOne row per package: version, description, author, homepage, repository, licence, keywords, page link, npm monthly downloads
Ceiling1,000 packages per npm search. Name lists are not capped
Account neededNone, and no registry key
Price$2.00 per 1,000 packages, flat on every plan. The free plan's $5 a month covers about 2,500

๐Ÿ” What Package Registry Scraper does

Two ways in, and you can use both at once on npm.

Exact names. Put a list in packageNames and each one is looked up directly. Scoped npm names like @types/node work. A name that does not exist gets its own NOT_FOUND row, uncharged, and the run carries on with the rest.

npm keyword search. Put words in searchQuery and the registry's own search decides what matches, ordered by its relevance score, up to your maxItems.

There is a real difference between the two that the field list hides: a search row has no licence, because npm's search index does not carry it. Repository, homepage and keywords come through on most search rows but not all. If you are doing a licence audit, feed the names in through packageNames and look them up properly.

includeDownloads is on by default and adds last month's download count to npm rows. PyPI has no public endpoint for that, so PyPI rows never carry it.

๐Ÿ“‹ What data you get from each npm or PyPI package

What you getField
Name, latest version and which registryname, version, registry
What it isdescription, keywords
Who wrote or published itauthor
Where its code and docs liverepository, homepage
The declared licencelicense
The project page on npmjs.com or pypi.orgurl
Last month's downloads, npm onlymonthlyDownloads
npm's search relevance, search rows onlyscore

โ–ถ๏ธ How to scrape npm and PyPI packages

  1. Open Package Registry Scraper and click Try for free.
  2. Pick a Registry.
  3. Paste names into Package names (exact), one per line. For an npm search, type into Search query (npm only) instead.
  4. Leave Include monthly downloads (npm) on unless you do not want it, then click Start.
  5. Download the dataset as JSON, CSV, Excel or XML.

๐Ÿ’ฐ How much does it cost to scrape npm and PyPI packages?

$2.00 per 1,000 packages. Flat on every Apify plan, no volume tiers, and the same on both registries. The free plan's $5 a month covers about 2,500 packages.

You pay per package row. The same name twice in one run is charged once. NOT_FOUND names, a search that matched nothing and every diagnostic row are not charged. Turning downloads on does not change the price, and a row whose download count could not be read is still a charged row.

If you set a maximum charge for a run, it stops at the last row that fits.

๐Ÿ“ฅ What you give it

{
"registry": "npm",
"packageNames": ["react", "@types/node"],
"includeDownloads": true
}
FieldDefaultWhat it is
registrynpmnpm or pypi.
packageNamesnoneExact names. Works on both registries, and is the only mode PyPI supports.
searchQuerynoneKeywords, npm only. Ignored on PyPI unless you sent no names, in which case you get BAD_INPUT. The Console shows react state management as an example, but that is a prefill, so an API call has to send its own.
maxItems501 to 1,000. Caps the npm search only. It does nothing to a list of names.
includeDownloadstrueLast month's downloads for npm packages. Costs one extra lookup per package, not extra money.
notionConnectornoneOptional. Writes every delivered package into your Notion. Authorise a connector once under Settings, API and Integrations, MCP connectors, then pick it here.
notionParentIdnoneOptional. The Notion data source id to write into. Leave it empty and the pages are created privately in your workspace.
proxyConfigurationoffOptional network setting. Off is right for a normal run.

packageNames has no ceiling. Paste a 4,000-line requirements.txt and you get up to 4,000 charged rows, because maxItems is not watching that path. Trim the list to what you need.

๐Ÿ“ค What you get back

A real row from a name lookup on 24 September 2026:

{
"ok": true,
"registry": "npm",
"name": "typescript",
"version": "7.0.2",
"description": "TypeScript is a language for application scale JavaScript development",
"author": "Microsoft Corp.",
"homepage": "https://www.typescriptlang.org/",
"repository": "https://github.com/microsoft/TypeScript",
"license": "Apache-2.0",
"keywords": ["TypeScript", "Microsoft", "compiler", "language", "javascript"],
"score": null,
"url": "https://www.npmjs.com/package/typescript",
"monthlyDownloads": 1004355023
}

And a real row from an npm search on 2 October 2026, which shows the search-path gap at its worst: no licence, and this package's repository, homepage and keywords are empty too.

{
"ok": true,
"registry": "npm",
"name": "unstated-next",
"version": "1.1.0",
"description": "200 bytes to never think about React state management libraries ever again",
"author": "thejameskyle",
"homepage": null,
"repository": null,
"license": null,
"keywords": [],
"score": 375.834,
"url": "https://www.npmjs.com/package/unstated-next",
"monthlyDownloads": 443321
}
FieldHow to read it
versionThe latest published version. npm's latest tag, PyPI's current release.
licenseThe declared licence string, as the package wrote it: Apache-2.0 on one, Apache License 2.0 on another. null on every npm search row.
repositoryNormalised into a clickable https URL rather than a git address.
scorenpm's own search relevance. Present on search rows, null on npm name lookups, and the key is absent from PyPI rows entirely.
monthlyDownloadsnpm only. null if the count could not be read, and the key is absent when includeDownloads is off or the row is from PyPI.
authorSometimes the publishing account's handle rather than a person or company: next comes back as vercel-release-bot.

๐Ÿงพ Reading the output

Packages carry ok: true. Anything with ok: false carries an errorCode and is not charged. One run can contain both, which is normal when a name list has a typo in it. There is no sample row.

CodeWhat it means
BAD_INPUTNeither names nor a query, or a PyPI run with a searchQuery and no names.
NOT_FOUNDThat one name is not in the registry. Check the spelling, the scope, and whether it was unpublished.
NO_RESULTSNothing at all came back. An npm search that matched no packages ends here.
NETWORKThe registry was unreachable or answered with something unusable. Re-run it.

The Console's Overview table hides license and repository, which are the two columns a licence audit actually needs. Download the dataset as JSON, CSV or Excel, or switch the table to all fields.

๐Ÿ’ก What people use it for

  • Auditing the licences behind a package.json or a requirements.txt by feeding the names in and reading the license column.
  • Comparing two libraries side by side on version, downloads and links before choosing one.
  • Watching a dependency on a schedule and firing an alert when version moves.
  • Sizing a niche: search npm for the keywords, then rank by monthlyDownloads to see what people actually install.

๐Ÿšง What it does not do

  • No PyPI search. Exact names only on that side.
  • No PyPI download counts. There is no public endpoint to read them from.
  • No licence on npm search rows. Use packageNames when the licence is the point.
  • No dependency trees and no version history. One row is the current state of one package.
  • No download trend. monthlyDownloads is last month's total, a single number, not a series.
  • No GitHub stars or issues. repository gives you the link, and GitHub Scraper reads the rest.
  • Rows are a snapshot. Versions and download counts move daily.
  • It is not an official npm tool. Dami's Studio is independent and is not affiliated with or endorsed by npm, PyPI or GitHub, or by any other company named on this page.

๐Ÿงญ Which developer data scraper do you need?

If you wantUse
npm and PyPI package metadata and downloadsThis one
Repositories, stars, forks and issues from GitHubGitHub Scraper
Stack Overflow questions, tags and scoresStack Overflow Scraper
Developer articles by tag or authorDEV Community Scraper
Show HN launches and developer discussionHacker News Scraper
Package lookups from inside Claude, Cursor or ChatGPT, beside other research toolsResearch MCP Server

โ“ Questions people ask

Do I need an npm or PyPI account or API key?

No. Both registries publish this openly, and the run needs nothing from you but the names or keywords.

Can I search PyPI by keyword?

No. PyPI has no clean public search API, so names are the only route. Find the name elsewhere and look it up here.

Why is the licence empty?

You came in through the npm search. Search results do not carry it. Take the names from that run and look them up in a second run.

Can I mix a search and a name list?

On npm, yes. Both run and the results are deduplicated.

Does maxItems protect me from a huge name list?

No. It caps the search only. Trim the list yourself.

Can I call it from code or connect it to an AI assistant?

Yes. The API tab has ready-made code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/package-registry-scraper. Either way the run happens on your Apify account at the same price.

Both registries publish this through open APIs, and it is package metadata rather than personal data, though author names can be people. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.

๐Ÿ†˜ If something breaks

Open the Issues tab on the actor page. Send the registry, the names or query you used and the run ID. The errorCode on the diagnostic row usually names the problem by itself.