# GitHub Repository Stats Scraper (`67-labs/github-repo-stats-scraper`) Actor

Get GitHub repo stats: stars, forks, open issues and PRs, contributors, releases, license, topics. CSV, JSON or Sheets.

- **URL**: https://apify.com/67-labs/github-repo-stats-scraper.md
- **Developed by:** [Mokksh Bhatt](https://apify.com/67-labs) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 result (one repository)s

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## GitHub Repository Stats Scraper — stars, forks, issues, pull requests, contributors and releases for any repo

Get the public numbers of any GitHub repository in one row: stars, forks, watchers, open issues, open pull requests, contributors, number of releases, the latest release, license, language, topics, and the dates of creation and last push. Give it `owner/name` or a GitHub link. Export to CSV, JSON, Excel or Google Sheets. It uses GitHub's official public API, so it is stable.

**Who uses this:** investors and analysts tracking open source projects, developer relations and marketing teams comparing tools, engineers choosing a library, researchers, recruiters mapping who builds what, AI agents that need current repository facts.

### What you get per repository

| Field | What it is |
|---|---|
| `repository`, `url` | Full name (`apify/crawlee`) and link |
| `description`, `homepage`, `language`, `topics`, `license` | The repository's own description, website, main language, topic tags and license (SPDX name) |
| `stars`, `forks`, `watchers` | The three public counters |
| `openIssuesAndPullRequests` | GitHub's own combined open counter |
| `openIssues`, `openPullRequests` | The same number split in two |
| `contributors` | People who contributed, including anonymous ones |
| `releasesCount`, `latestReleaseTag`, `latestReleaseDate`, `latestReleaseUrl`, `latestReleaseIsPrerelease` | How many releases exist and the newest one |
| `createdAt`, `pushedAt` | When the repository was created and when code was last pushed |
| `sizeKb`, `defaultBranch`, `archived`, `isFork`, `ownerType` | Size, main branch, and flags |
| `scrapedAt` | When the row was collected |

A field is `null` when GitHub does not provide it, for example the pull request numbers of a repository that does not take pull requests.

### Price

**$0.01 per run + $2 per 1,000 repositories.** That is two pay-per-event charges: `actor-start` ($0.01) once per run and `result` ($0.002) per repository. No compute fees on top. The free Apify plan ($5 per month credit) covers about 2,400 repositories a month at no cost. A repository that is not found is not charged. A run that fails before it reads anything is not charged.

### Use a GitHub token for more than about 15 repositories

GitHub allows **60 requests per hour** without a token, and one repository needs 4 requests, so about 15 repositories an hour. Apify runs share addresses, so the limit can be used up by other users. A free token raises the limit to **5,000 requests an hour**, about 1,200 repositories:

1. In GitHub open Settings, Developer settings, Personal access tokens, and create a token. It needs no permissions for public repositories.
2. Paste it in the **GitHub token** field. It is stored encrypted and only sent to `api.github.com`.

With **Include contributors, releases and pull requests** turned off, one repository needs only 1 request.

### How to use

1. Open the Input tab and add repositories, one per line: `owner/name` or the GitHub link.
2. Paste your GitHub token if you have more than a few repositories.
3. Click Start. Open the Output tab and export the table as CSV, JSON or Excel, or send it to Google Sheets with Apify's Google Sheets integration.

### Input example

```json
{
    "repositories": ["apify/crawlee", "https://github.com/microsoft/playwright"],
    "includeDetails": true
}
```

### Output example

```json
{
    "repository": "apify/crawlee",
    "url": "https://github.com/apify/crawlee",
    "description": "Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. ...",
    "homepage": "https://crawlee.dev",
    "language": "TypeScript",
    "topics": ["apify", "automation", "crawler", "playwright", "puppeteer", "scraping", "typescript"],
    "license": "Apache-2.0",
    "stars": 25900,
    "forks": 1681,
    "watchers": 136,
    "openIssuesAndPullRequests": 135,
    "openIssues": 95,
    "openPullRequests": 40,
    "contributors": 148,
    "releasesCount": 137,
    "latestReleaseTag": "v4.0.0-rc.0",
    "latestReleaseDate": "2026-08-13T10:14:22Z",
    "latestReleaseUrl": "https://github.com/apify/crawlee/releases/tag/v4.0.0-rc.0",
    "latestReleaseIsPrerelease": true,
    "createdAt": "2016-08-26T18:35:03Z",
    "pushedAt": "2026-09-25T21:28:50Z",
    "sizeKb": 175650,
    "defaultBranch": "master",
    "archived": false,
    "isFork": false,
    "ownerType": "Organization",
    "scrapedAt": "2026-09-26T10:00:00.000Z"
}
```

The topics list is shortened in this example.

### Limits

- Public repositories only. A private repository looks like "not found".
- Very large repositories can be missing some numbers: GitHub does not return contributors for some very large repositories, and some repositories take no pull requests. Those fields are `null`. The other fields are still returned.
- `latestRelease` is the newest release GitHub lists, including pre-releases. Check `latestReleaseIsPrerelease`.
- If GitHub's rate limit is reached, the run stops reading, saves what it has, and the log says when the limit resets.

### FAQ

**Can I track stars over time?** Yes. Schedule the actor in Apify with the same repositories and compare `stars` between runs.

**Does it need a proxy?** No. GitHub limits by token or address, so a proxy only helps a run without a token.

**Can an AI agent use it?** Yes. It works through Apify's API and MCP server. Give it a list of repositories.

**Something is wrong. What now?** Open an issue on the Issues tab with the repository and the input you used. Field requests are welcome.

### Changelog

- **0.1** First release: repository stats, contributors, releases, pull requests, optional token, pay per event.

# Actor input Schema

## `repositories` (type: `array`):

One entry per repository: owner/name (apify/crawlee) or the GitHub link. You pay per repository, so this list is also your cost cap.

## `githubToken` (type: `string`):

Without a token GitHub allows only 60 requests per hour, about 15 repositories. A free token (a personal access token with no permissions) raises it to 5,000 per hour. The token is stored encrypted and only sent to api.github.com.

## `includeDetails` (type: `boolean`):

Adds the contributor count, the release count and the open pull requests. Needs 3 more requests per repository. Turn off for a fast list of stars, forks and issues.

## `proxyConfiguration` (type: `object`):

Leave off. GitHub limits by token or address, so a proxy only helps a run without a token.

## Actor input object example

```json
{
  "repositories": [
    "apify/crawlee",
    "https://github.com/microsoft/playwright"
  ],
  "includeDetails": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `repositories` (type: `string`):

One row per repository: stars, forks, issues, pull requests, contributors, releases, license, topics.

## `repositoriesCsv` (type: `string`):

Same rows as CSV for Google Sheets or Excel.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "repositories": [
        "apify/crawlee"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("67-labs/github-repo-stats-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "repositories": ["apify/crawlee"] }

# Run the Actor and wait for it to finish
run = client.actor("67-labs/github-repo-stats-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "repositories": [
    "apify/crawlee"
  ]
}' |
apify call 67-labs/github-repo-stats-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,67-labs/github-repo-stats-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/LhuG04dXfdaX4cKkL/builds/LnhwhOY44zvGjemoH/openapi.json
