# PyPI Scraper - Package Metadata, Deps and Releases (`s-r/pypi-scraper`) Actor

Pull PyPI package metadata by name or name pattern: version, licence, declared dependencies, supported Python versions, full release history with dates, and yanked releases. Reads PyPI's own JSON API and the simple index, no key needed.

- **URL**: https://apify.com/s-r/pypi-scraper.md
- **Developed by:** [SR](https://apify.com/s-r) (community)
- **Categories:** Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 run start fees

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PyPI Scraper

Pull metadata for Python packages straight from **PyPI's own JSON API**: current
version, licence, declared dependencies, supported Python versions, the full
release history with dates, and which releases were **yanked** after
publication. No key, no login, no scraping of the website.

Give it a list of package names, or a name pattern to match across the whole
index of **883,837 projects**.

### What this answers that a package page does not

**"What does this actually depend on?"** `requires_dist` is the real dependency
list with version specifiers and environment markers, exactly as the maintainer
declared it. That is what a supply-chain review or a licence audit needs, and
reading it off the rendered page is guesswork.

**"Is this project still alive?"** Every row carries `last_release_at` and
`days_since_last_release`. Sort by it and an unmaintained dependency stands out
immediately. The run summary counts how many of your packages have had no
release in over two years.

**"Are we pinned to a release that was pulled?"** `yanked_versions` lists every
version whose files were all yanked. Maintainers yank releases for broken
builds, licence mistakes and security problems, and a pinned dependency on one
is a finding that appears nowhere in a naive scrape. A version with even one
live file is *not* counted, because it is still installable.

### Two ways to use it

**By name**, the fast path. One API call per package, ten at a time:

```
packages: ["httpx", "requests", "django"]
```

**By pattern**, when you do not know the names:

```
name_pattern: "django-rest*"
```

A pattern with no wildcard is treated as "contains", so `boto` matches every
project with `boto` in the name. Matches are returned shortest-name-first, which
puts the canonical project ahead of its forks and plugins.

### About the pattern search

**PyPI publishes no search API.** The `/search/` page is HTML, hands about 3 KB
to anything that is not a browser, and the maintainers ask people not to scrape
it. So this Actor does not.

Instead a pattern is matched against PyPI's **simple index**, the official
machine-readable list of every project, requested with the documented JSON
accept header. It is one download of the full index followed by local filtering.
That costs a few seconds more at the start of a run and asks a service funded by
donations for exactly one page instead of many.

If you already know the names, pass `packages` and skip the index entirely.

### Fields

- **Identity**: `name`, `version`, `summary`, `keywords`, `url`, `package_url`
- **Legal**: `license`, `license_classifiers` (the trove classifiers, which are
  often more reliable than the free-text field)
- **People**: `author`, `author_email`, `maintainer`
- **Links**: `home_page`, `project_urls` (docs, changelog, source, issues)
- **Compatibility**: `requires_python`, `python_versions`,
  `development_status`, `classifiers`
- **Dependencies**: `requires_dist`, `dependency_count`
- **History**: `release_count`, `first_release_at`, `last_release_at`,
  `days_since_last_release`
- **Yanking**: `yanked`, `yanked_reason`, `yanked_versions`
- **Current release files**: `latest_files`, `latest_size_bytes`, `has_wheel`

`has_wheel` is worth a note: a project shipping only a source distribution has
to compile on install, which is the difference between a fast CI run and a slow
one that needs a toolchain.

### A missing package is said, not implied

A name that does not exist on PyPI returns a `not_found` error naming it, rather
than a row of nulls. When you are auditing a requirements file, "this package
does not exist" and "this package exists but has no dependencies" are entirely
different findings, and a scraper that blurs them is worse than useless.

Transport failures are separated too. A 404 is definitive and reported as
`not_found`; a 500 or a timeout is retried with backoff and only then reported
as `fetch_failed`. You always know which of the two you are looking at.

### Input reference

| Field | Type | Default |
|---|---|---|
| `packages` | list of project names | `["httpx", "requests", "django"]` |
| `name_pattern` | glob or substring | — |
| `limit` | 1-1000 | 50 |
| `retries` | 1-6 | 3 |

Give either a list of names or a pattern. An empty input is rejected with a
message rather than walking the index for no reason.

### Typical uses

- **Dependency and licence audit.** Feed your requirements file in, get every
  declared licence and dependency out, then filter on the licences your legal
  team cares about.
- **Maintenance review.** Sort your dependency tree by
  `days_since_last_release` and see what has been abandoned under you.
- **Yanked-release check.** Cross-reference your pinned versions against
  `yanked_versions`.
- **Ecosystem research.** Match a pattern like `django-*` or `*-airflow-*` and
  measure how a plugin ecosystem is doing: how many projects, how many still
  releasing, which Python versions they have moved to.

### Notes

PyPI metadata is only as good as what maintainers declare. `license` is a free
text field and is frequently empty even when `license_classifiers` is populated,
which is why both are returned. `author` and `home_page` are increasingly left
blank in favour of `project_urls`, so that is checked as a fallback for the
homepage.

# Actor input Schema

## `packages` (type: `array`):

Exact PyPI project names to look up, for example httpx or django. This is the fast path: each name is one API call.

## `name_pattern` (type: `string`):

Match project names instead of listing them, for example django-rest\* or anything containing 'boto'. PyPI has no search API, so this downloads the full project index once and filters locally. Slower to start, and much kinder to PyPI than scraping the search page.

## `limit` (type: `integer`):

How many projects to return.

## `retries` (type: `integer`):

Retries with backoff before a request is reported as an error.

## Actor input object example

```json
{
  "packages": [
    "httpx",
    "requests",
    "django"
  ],
  "name_pattern": "django-rest*",
  "limit": 50,
  "retries": 3
}
```

# Actor output Schema

## `packages` (type: `string`):

One row per PyPI project.

## `summary` (type: `string`):

Counts, dependency coverage and how many projects look unmaintained.

## `errors` (type: `string`):

Failures with a code and a redacted message.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "packages": [
        "httpx",
        "requests",
        "django"
    ],
    "limit": 50,
    "retries": 3
};

// Run the Actor and wait for it to finish
const run = await client.actor("s-r/pypi-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "packages": [
        "httpx",
        "requests",
        "django",
    ],
    "limit": 50,
    "retries": 3,
}

# Run the Actor and wait for it to finish
run = client.actor("s-r/pypi-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "packages": [
    "httpx",
    "requests",
    "django"
  ],
  "limit": 50,
  "retries": 3
}' |
apify call s-r/pypi-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,s-r/pypi-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WjxfM1i0fgzYnV5Gt/builds/Z5uFjsfC9pvhKlcjm/openapi.json
