npm + PyPI Package Scraper - Metadata & Downloads
Pricing
from $1.94 / 1,000 packages
npm + PyPI Package Scraper - Metadata & Downloads
Look up npm and PyPI packages by exact name, or search npm by keyword. Each row has the latest version, description, author, repository, licence, keywords and project link, plus monthly downloads on npm. No API key. PyPI matches exact names only. $2 per 1,000 packages.
Pricing
from $1.94 / 1,000 packages
Rating
0.0
(0)
Developer
Dami's Studio
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
1
Monthly active users
21 hours ago
Last modified
Categories
Share
Look up npm and PyPI packages by exact name, or search npm by keyword, and get one row per package: the latest version, description, author, homepage, repository, licence, keywords and the project page, plus last month's download count on npm. Both registries come back in the same row shape.
The two registries are not symmetrical, and it matters before you plan a run. npm does keyword search and exact names. PyPI does exact names only, because it publishes no clean search API for anyone to call. And an npm search row never carries the licence, so a licence audit goes through names.
| Input | Exact package names on npm or PyPI, or keywords for an npm search |
| Output | One row per package: version, description, author, homepage, repository, licence, keywords, page link, npm monthly downloads |
| Ceiling | 1,000 packages per npm search. Name lists are not capped |
| Account needed | None, and no registry key |
| Price | $2.00 per 1,000 packages, flat on every plan. The free plan's $5 a month covers about 2,500 |
๐ What Package Registry Scraper does
Two ways in, and you can use both at once on npm.
Exact names. Put a list in packageNames and each one is looked up directly. Scoped npm names like @types/node work. A name that does not exist gets its own NOT_FOUND row, uncharged, and the run carries on with the rest.
npm keyword search. Put words in searchQuery and the registry's own search decides what matches, ordered by its relevance score, up to your maxItems.
There is a real difference between the two that the field list hides: a search row has no licence, because npm's search index does not carry it. Repository, homepage and keywords come through on most search rows but not all. If you are doing a licence audit, feed the names in through packageNames and look them up properly.
includeDownloads is on by default and adds last month's download count to npm rows. PyPI has no public endpoint for that, so PyPI rows never carry it.
๐ What data you get from each npm or PyPI package
| What you get | Field |
|---|---|
| Name, latest version and which registry | name, version, registry |
| What it is | description, keywords |
| Who wrote or published it | author |
| Where its code and docs live | repository, homepage |
| The declared licence | license |
| The project page on npmjs.com or pypi.org | url |
| Last month's downloads, npm only | monthlyDownloads |
| npm's search relevance, search rows only | score |
โถ๏ธ How to scrape npm and PyPI packages
- Open Package Registry Scraper and click Try for free.
- Pick a Registry.
- Paste names into Package names (exact), one per line. For an npm search, type into Search query (npm only) instead.
- Leave Include monthly downloads (npm) on unless you do not want it, then click Start.
- Download the dataset as JSON, CSV, Excel or XML.
๐ฐ How much does it cost to scrape npm and PyPI packages?
$2.00 per 1,000 packages. Flat on every Apify plan, no volume tiers, and the same on both registries. The free plan's $5 a month covers about 2,500 packages.
You pay per package row. The same name twice in one run is charged once. NOT_FOUND names, a search that matched nothing and every diagnostic row are not charged. Turning downloads on does not change the price, and a row whose download count could not be read is still a charged row.
If you set a maximum charge for a run, it stops at the last row that fits.
๐ฅ What you give it
{"registry": "npm","packageNames": ["react", "@types/node"],"includeDownloads": true}
| Field | Default | What it is |
|---|---|---|
registry | npm | npm or pypi. |
packageNames | none | Exact names. Works on both registries, and is the only mode PyPI supports. |
searchQuery | none | Keywords, npm only. Ignored on PyPI unless you sent no names, in which case you get BAD_INPUT. The Console shows react state management as an example, but that is a prefill, so an API call has to send its own. |
maxItems | 50 | 1 to 1,000. Caps the npm search only. It does nothing to a list of names. |
includeDownloads | true | Last month's downloads for npm packages. Costs one extra lookup per package, not extra money. |
notionConnector | none | Optional. Writes every delivered package into your Notion. Authorise a connector once under Settings, API and Integrations, MCP connectors, then pick it here. |
notionParentId | none | Optional. The Notion data source id to write into. Leave it empty and the pages are created privately in your workspace. |
proxyConfiguration | off | Optional network setting. Off is right for a normal run. |
packageNames has no ceiling. Paste a 4,000-line requirements.txt and you get up to 4,000 charged rows, because maxItems is not watching that path. Trim the list to what you need.
๐ค What you get back
A real row from a name lookup on 24 September 2026:
{"ok": true,"registry": "npm","name": "typescript","version": "7.0.2","description": "TypeScript is a language for application scale JavaScript development","author": "Microsoft Corp.","homepage": "https://www.typescriptlang.org/","repository": "https://github.com/microsoft/TypeScript","license": "Apache-2.0","keywords": ["TypeScript", "Microsoft", "compiler", "language", "javascript"],"score": null,"url": "https://www.npmjs.com/package/typescript","monthlyDownloads": 1004355023}
And a real row from an npm search on 2 October 2026, which shows the search-path gap at its worst: no licence, and this package's repository, homepage and keywords are empty too.
{"ok": true,"registry": "npm","name": "unstated-next","version": "1.1.0","description": "200 bytes to never think about React state management libraries ever again","author": "thejameskyle","homepage": null,"repository": null,"license": null,"keywords": [],"score": 375.834,"url": "https://www.npmjs.com/package/unstated-next","monthlyDownloads": 443321}
| Field | How to read it |
|---|---|
version | The latest published version. npm's latest tag, PyPI's current release. |
license | The declared licence string, as the package wrote it: Apache-2.0 on one, Apache License 2.0 on another. null on every npm search row. |
repository | Normalised into a clickable https URL rather than a git address. |
score | npm's own search relevance. Present on search rows, null on npm name lookups, and the key is absent from PyPI rows entirely. |
monthlyDownloads | npm only. null if the count could not be read, and the key is absent when includeDownloads is off or the row is from PyPI. |
author | Sometimes the publishing account's handle rather than a person or company: next comes back as vercel-release-bot. |
๐งพ Reading the output
Packages carry ok: true. Anything with ok: false carries an errorCode and is not charged. One run can contain both, which is normal when a name list has a typo in it. There is no sample row.
| Code | What it means |
|---|---|
BAD_INPUT | Neither names nor a query, or a PyPI run with a searchQuery and no names. |
NOT_FOUND | That one name is not in the registry. Check the spelling, the scope, and whether it was unpublished. |
NO_RESULTS | Nothing at all came back. An npm search that matched no packages ends here. |
NETWORK | The registry was unreachable or answered with something unusable. Re-run it. |
The Console's Overview table hides license and repository, which are the two columns a licence audit actually needs. Download the dataset as JSON, CSV or Excel, or switch the table to all fields.
๐ก What people use it for
- Auditing the licences behind a
package.jsonor arequirements.txtby feeding the names in and reading thelicensecolumn. - Comparing two libraries side by side on version, downloads and links before choosing one.
- Watching a dependency on a schedule and firing an alert when
versionmoves. - Sizing a niche: search npm for the keywords, then rank by
monthlyDownloadsto see what people actually install.
๐ง What it does not do
- No PyPI search. Exact names only on that side.
- No PyPI download counts. There is no public endpoint to read them from.
- No licence on npm search rows. Use
packageNameswhen the licence is the point. - No dependency trees and no version history. One row is the current state of one package.
- No download trend.
monthlyDownloadsis last month's total, a single number, not a series. - No GitHub stars or issues.
repositorygives you the link, and GitHub Scraper reads the rest. - Rows are a snapshot. Versions and download counts move daily.
- It is not an official npm tool. Dami's Studio is independent and is not affiliated with or endorsed by npm, PyPI or GitHub, or by any other company named on this page.
๐งญ Which developer data scraper do you need?
| If you want | Use |
|---|---|
| npm and PyPI package metadata and downloads | This one |
| Repositories, stars, forks and issues from GitHub | GitHub Scraper |
| Stack Overflow questions, tags and scores | Stack Overflow Scraper |
| Developer articles by tag or author | DEV Community Scraper |
| Show HN launches and developer discussion | Hacker News Scraper |
| Package lookups from inside Claude, Cursor or ChatGPT, beside other research tools | Research MCP Server |
โ Questions people ask
Do I need an npm or PyPI account or API key?
No. Both registries publish this openly, and the run needs nothing from you but the names or keywords.
Can I search PyPI by keyword?
No. PyPI has no clean public search API, so names are the only route. Find the name elsewhere and look it up here.
Why is the licence empty?
You came in through the npm search. Search results do not carry it. Take the names from that run and look them up in a second run.
Can I mix a search and a name list?
On npm, yes. Both run and the results are deduplicated.
Does maxItems protect me from a huge name list?
No. It caps the search only. Trim the list yourself.
Can I call it from code or connect it to an AI assistant?
Yes. The API tab has ready-made code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/package-registry-scraper. Either way the run happens on your Apify account at the same price.
Is it legal to scrape npm and PyPI?
Both registries publish this through open APIs, and it is package metadata rather than personal data, though author names can be people. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.
๐ If something breaks
Open the Issues tab on the actor page. Send the registry, the names or query you used and the run ID. The errorCode on the diagnostic row usually names the problem by itself.