PyPI Scraper - Package Metadata, Deps and Releases
Pricing
from $1.00 / 1,000 run start fees
PyPI Scraper - Package Metadata, Deps and Releases
Pull PyPI package metadata by name or name pattern: version, licence, declared dependencies, supported Python versions, full release history with dates, and yanked releases. Reads PyPI's own JSON API and the simple index, no key needed.
Pricing
from $1.00 / 1,000 run start fees
Rating
0.0
(0)
Developer
SR
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
PyPI Scraper
Pull metadata for Python packages straight from PyPI's own JSON API: current version, licence, declared dependencies, supported Python versions, the full release history with dates, and which releases were yanked after publication. No key, no login, no scraping of the website.
Give it a list of package names, or a name pattern to match across the whole index of 883,837 projects.
What this answers that a package page does not
"What does this actually depend on?" requires_dist is the real dependency
list with version specifiers and environment markers, exactly as the maintainer
declared it. That is what a supply-chain review or a licence audit needs, and
reading it off the rendered page is guesswork.
"Is this project still alive?" Every row carries last_release_at and
days_since_last_release. Sort by it and an unmaintained dependency stands out
immediately. The run summary counts how many of your packages have had no
release in over two years.
"Are we pinned to a release that was pulled?" yanked_versions lists every
version whose files were all yanked. Maintainers yank releases for broken
builds, licence mistakes and security problems, and a pinned dependency on one
is a finding that appears nowhere in a naive scrape. A version with even one
live file is not counted, because it is still installable.
Two ways to use it
By name, the fast path. One API call per package, ten at a time:
packages: ["httpx", "requests", "django"]
By pattern, when you do not know the names:
name_pattern: "django-rest*"
A pattern with no wildcard is treated as "contains", so boto matches every
project with boto in the name. Matches are returned shortest-name-first, which
puts the canonical project ahead of its forks and plugins.
About the pattern search
PyPI publishes no search API. The /search/ page is HTML, hands about 3 KB
to anything that is not a browser, and the maintainers ask people not to scrape
it. So this Actor does not.
Instead a pattern is matched against PyPI's simple index, the official machine-readable list of every project, requested with the documented JSON accept header. It is one download of the full index followed by local filtering. That costs a few seconds more at the start of a run and asks a service funded by donations for exactly one page instead of many.
If you already know the names, pass packages and skip the index entirely.
Fields
- Identity:
name,version,summary,keywords,url,package_url - Legal:
license,license_classifiers(the trove classifiers, which are often more reliable than the free-text field) - People:
author,author_email,maintainer - Links:
home_page,project_urls(docs, changelog, source, issues) - Compatibility:
requires_python,python_versions,development_status,classifiers - Dependencies:
requires_dist,dependency_count - History:
release_count,first_release_at,last_release_at,days_since_last_release - Yanking:
yanked,yanked_reason,yanked_versions - Current release files:
latest_files,latest_size_bytes,has_wheel
has_wheel is worth a note: a project shipping only a source distribution has
to compile on install, which is the difference between a fast CI run and a slow
one that needs a toolchain.
A missing package is said, not implied
A name that does not exist on PyPI returns a not_found error naming it, rather
than a row of nulls. When you are auditing a requirements file, "this package
does not exist" and "this package exists but has no dependencies" are entirely
different findings, and a scraper that blurs them is worse than useless.
Transport failures are separated too. A 404 is definitive and reported as
not_found; a 500 or a timeout is retried with backoff and only then reported
as fetch_failed. You always know which of the two you are looking at.
Input reference
| Field | Type | Default |
|---|---|---|
packages | list of project names | ["httpx", "requests", "django"] |
name_pattern | glob or substring | — |
limit | 1-1000 | 50 |
retries | 1-6 | 3 |
Give either a list of names or a pattern. An empty input is rejected with a message rather than walking the index for no reason.
Typical uses
- Dependency and licence audit. Feed your requirements file in, get every declared licence and dependency out, then filter on the licences your legal team cares about.
- Maintenance review. Sort your dependency tree by
days_since_last_releaseand see what has been abandoned under you. - Yanked-release check. Cross-reference your pinned versions against
yanked_versions. - Ecosystem research. Match a pattern like
django-*or*-airflow-*and measure how a plugin ecosystem is doing: how many projects, how many still releasing, which Python versions they have moved to.
Notes
PyPI metadata is only as good as what maintainers declare. license is a free
text field and is frequently empty even when license_classifiers is populated,
which is why both are returned. author and home_page are increasingly left
blank in favour of project_urls, so that is checked as a fallback for the
homepage.