PyPI Scraper - Package Metadata, Deps and Releases avatar

PyPI Scraper - Package Metadata, Deps and Releases

Pricing

from $1.00 / 1,000 run start fees

Go to Apify Store
PyPI Scraper - Package Metadata, Deps and Releases

PyPI Scraper - Package Metadata, Deps and Releases

Pull PyPI package metadata by name or name pattern: version, licence, declared dependencies, supported Python versions, full release history with dates, and yanked releases. Reads PyPI's own JSON API and the simple index, no key needed.

Pricing

from $1.00 / 1,000 run start fees

Rating

0.0

(0)

Developer

SR

SR

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Categories

Share

PyPI Scraper

Pull metadata for Python packages straight from PyPI's own JSON API: current version, licence, declared dependencies, supported Python versions, the full release history with dates, and which releases were yanked after publication. No key, no login, no scraping of the website.

Give it a list of package names, or a name pattern to match across the whole index of 883,837 projects.

What this answers that a package page does not

"What does this actually depend on?" requires_dist is the real dependency list with version specifiers and environment markers, exactly as the maintainer declared it. That is what a supply-chain review or a licence audit needs, and reading it off the rendered page is guesswork.

"Is this project still alive?" Every row carries last_release_at and days_since_last_release. Sort by it and an unmaintained dependency stands out immediately. The run summary counts how many of your packages have had no release in over two years.

"Are we pinned to a release that was pulled?" yanked_versions lists every version whose files were all yanked. Maintainers yank releases for broken builds, licence mistakes and security problems, and a pinned dependency on one is a finding that appears nowhere in a naive scrape. A version with even one live file is not counted, because it is still installable.

Two ways to use it

By name, the fast path. One API call per package, ten at a time:

packages: ["httpx", "requests", "django"]

By pattern, when you do not know the names:

name_pattern: "django-rest*"

A pattern with no wildcard is treated as "contains", so boto matches every project with boto in the name. Matches are returned shortest-name-first, which puts the canonical project ahead of its forks and plugins.

PyPI publishes no search API. The /search/ page is HTML, hands about 3 KB to anything that is not a browser, and the maintainers ask people not to scrape it. So this Actor does not.

Instead a pattern is matched against PyPI's simple index, the official machine-readable list of every project, requested with the documented JSON accept header. It is one download of the full index followed by local filtering. That costs a few seconds more at the start of a run and asks a service funded by donations for exactly one page instead of many.

If you already know the names, pass packages and skip the index entirely.

Fields

  • Identity: name, version, summary, keywords, url, package_url
  • Legal: license, license_classifiers (the trove classifiers, which are often more reliable than the free-text field)
  • People: author, author_email, maintainer
  • Links: home_page, project_urls (docs, changelog, source, issues)
  • Compatibility: requires_python, python_versions, development_status, classifiers
  • Dependencies: requires_dist, dependency_count
  • History: release_count, first_release_at, last_release_at, days_since_last_release
  • Yanking: yanked, yanked_reason, yanked_versions
  • Current release files: latest_files, latest_size_bytes, has_wheel

has_wheel is worth a note: a project shipping only a source distribution has to compile on install, which is the difference between a fast CI run and a slow one that needs a toolchain.

A missing package is said, not implied

A name that does not exist on PyPI returns a not_found error naming it, rather than a row of nulls. When you are auditing a requirements file, "this package does not exist" and "this package exists but has no dependencies" are entirely different findings, and a scraper that blurs them is worse than useless.

Transport failures are separated too. A 404 is definitive and reported as not_found; a 500 or a timeout is retried with backoff and only then reported as fetch_failed. You always know which of the two you are looking at.

Input reference

FieldTypeDefault
packageslist of project names["httpx", "requests", "django"]
name_patternglob or substring—
limit1-100050
retries1-63

Give either a list of names or a pattern. An empty input is rejected with a message rather than walking the index for no reason.

Typical uses

  • Dependency and licence audit. Feed your requirements file in, get every declared licence and dependency out, then filter on the licences your legal team cares about.
  • Maintenance review. Sort your dependency tree by days_since_last_release and see what has been abandoned under you.
  • Yanked-release check. Cross-reference your pinned versions against yanked_versions.
  • Ecosystem research. Match a pattern like django-* or *-airflow-* and measure how a plugin ecosystem is doing: how many projects, how many still releasing, which Python versions they have moved to.

Notes

PyPI metadata is only as good as what maintainers declare. license is a free text field and is frequently empty even when license_classifiers is populated, which is why both are returned. author and home_page are increasingly left blank in favour of project_urls, so that is checked as a fallback for the homepage.