# GitHub Repo Scraper — Stars, Forks, Topics & Activity (`hichemdev/github-repo-scraper`) Actor

Scrape GitHub repositories with stars, forks, watchers, open issues, language, topics, licence, size and created, updated and last-push dates. Search by language, topic, stars, owner or any GitHub query, list an organisation's repos, or fetch exact repositories. No API key.

- **URL**: https://apify.com/hichemdev/github-repo-scraper.md
- **Developed by:** [Hichem Ben Moussa](https://apify.com/hichemdev) (community)
- **Categories:** Developer tools, Business, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 repositories

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## GitHub Repo Scraper — Stars, Forks, Topics & Activity

Scrape **GitHub repositories with their full metadata** — stars, forks, watchers, open issues, language, topics, licence, size and the created, updated and last-push dates — by search query, by language and star count, by organisation, or by exact repository name.

No API key required for small runs; add a free token for large ones.

### What you get

| Field | Description |
|---|---|
| `fullName`, `name`, `owner`, `ownerType` | Identity, and whether the owner is a User or Organization |
| `description` | Repository description |
| `stars`, `forks`, `watchers` | Popularity |
| `openIssues` | Open issues and pull requests |
| `language` | Primary language |
| `topics` | Topic tags |
| `license` | SPDX identifier, e.g. `MIT`, `Apache-2.0` |
| `homepage` | Project site |
| `isFork`, `isArchived`, `isTemplate` | Repository status |
| `hasIssues`, `hasWiki`, `hasDiscussions` | Enabled features |
| `defaultBranch`, `sizeKb` | Branch and checkout size |
| `createdAt`, `updatedAt`, `pushedAt` | **`pushedAt` is the real activity signal** — `updatedAt` also moves on metadata edits |
| `url`, `cloneUrl`, `apiUrl` | Links |

### Example input

```json
{
  "language": "rust",
  "minStars": 10000,
  "sortBy": "stars",
  "maxRepos": 100
}
```

Find actively maintained projects in a niche:

```json
{
  "topic": "kubernetes",
  "minStars": 500,
  "pushedAfter": "2026-06-01"
}
```

List everything an organisation owns, or fetch exact repositories:

```json
{ "owner": "vercel", "maxRepos": 500 }
{ "repos": ["facebook/react", "torvalds/linux"] }
```

Power users can pass any GitHub search string directly:

```json
{ "searchQuery": "language:go stars:>5000 topic:cli archived:false" }
```

### Example output

```json
{
  "fullName": "clash-verge-rev/clash-verge-rev",
  "owner": "clash-verge-rev",
  "ownerType": "Organization",
  "description": "A modern GUI client based on Tauri...",
  "stars": 147543,
  "forks": 10600,
  "openIssues": 474,
  "language": "Rust",
  "topics": ["clash", "clash-meta", "clash-verge", "linux"],
  "license": "GPL-3.0",
  "pushedAt": "2026-09-26T10:01:51Z",
  "isArchived": false,
  "url": "https://github.com/clash-verge-rev/clash-verge-rev"
}
```

### Who uses this

- **Developer-tool marketing** — find the repositories, and therefore the maintainers and communities, in your category
- **Technical recruiting** — surface active projects in a language or framework, then their owners
- **Open-source strategy** — benchmark your project against competitors on stars, issue load and push cadence
- **VC and deal sourcing** — fast-growing new repos are a leading indicator; combine `createdAfter` with `minStars`
- **Security and supply chain** — inventory an org's repos with licences, archived status and last-push dates
- **Dependency and licence audits** — `license` plus `isArchived` finds abandoned dependencies before they bite

### Rate limits — read this for large runs

GitHub allows **60 requests per hour without authentication** and **5,000 with a token**. Each request returns up to 100 repositories, so an unauthenticated run comfortably handles a few thousand repos per hour, and the actor reports the remaining quota in the log.

For anything bigger, create a free personal access token at github.com → Settings → Developer settings → Personal access tokens and paste it into the **GitHub token** field. **No scopes need to be granted** for public repository data. If you do hit the limit without a token, the actor says so explicitly rather than returning a silent partial result.

### Two API limits worth knowing

- **GitHub's search API never returns more than 1,000 results per query**, no matter the page. The actor reports the true total match count so you can see when you are hitting it, and the fix is to narrow the query — split by star band (`stars:1000..5000`), by language, or by creation year.
- **Listing by owner alone bypasses that cap.** When you set only `owner`, the actor uses the repository-listing endpoint rather than search, so it can walk an organisation's entire catalogue.

### Pricing

Pay per result. Each repository returned counts as one result.

### Notes

- `watchers` is GitHub's true subscriber count where available; the search API returns a field of the same name that actually mirrors the star count, and the actor prefers the real one.
- `license` is the SPDX identifier; repositories with a non-standard licence file report the licence name instead, and unlicensed repositories report null.
- This is an unofficial actor and is not affiliated with GitHub.

# Actor input Schema

## `searchQuery` (type: `string`):

Advanced. A raw GitHub search string, e.g. language:go stars:>5000 topic:cli. Overrides the simple filters below.

## `language` (type: `string`):

Primary language, e.g. python, rust, typescript.

## `topic` (type: `string`):

GitHub topic tag, e.g. machine-learning, kubernetes.

## `minStars` (type: `integer`):

Keep only repositories with at least this many stars.

## `owner` (type: `string`):

A user or org login. On its own it lists all of their repositories, bypassing the search API's 1,000-result cap.

## `repos` (type: `array`):

Fetch these exact repositories instead of searching, as owner/name or a GitHub URL, e.g. facebook/react.

## `pushedAfter` (type: `string`):

YYYY-MM-DD. Useful for filtering out abandoned projects.

## `createdAfter` (type: `string`):

YYYY-MM-DD. Useful for spotting fast-growing new projects.

## `sortBy` (type: `string`):

Result order for searches.

## `githubToken` (type: `string`):

Without a token GitHub allows 60 requests an hour; with a free personal access token it allows 5,000. Needed for large runs. No scopes are required for public repositories.

## `maxRepos` (type: `integer`):

Stop after this many repositories.

## `proxyConfiguration` (type: `object`):

Optional. A proxy is not required for this actor.

## Actor input object example

```json
{
  "language": "rust",
  "minStars": 10000,
  "sortBy": "stars",
  "maxRepos": 50,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

All scraped items as JSON.

## `overview` (type: `string`):

Browse results in the Apify Console.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "language": "rust",
    "minStars": 10000,
    "maxRepos": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("hichemdev/github-repo-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "language": "rust",
    "minStars": 10000,
    "maxRepos": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("hichemdev/github-repo-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "language": "rust",
  "minStars": 10000,
  "maxRepos": 50
}' |
apify call hichemdev/github-repo-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,hichemdev/github-repo-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/z3jmZBPiqNlmxeyDW/builds/dqUcvU2JI7dTdAIWk/openapi.json
