# GitHub Repo Scraper – Trending, Search, Releases & Issues API (`gazidev/github-repo-data`) Actor

Get GitHub repository data from the official REST API: trending new repos by stars, search with filters, all repos of a user or org, or details for a list. Stars, forks, topics, license, languages, contributors count, releases and issues/PRs. Works without a token; no emails collected.

- **URL**: https://apify.com/gazidev/github-repo-data.md
- **Developed by:** [Cemal Atakli](https://apify.com/gazidev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.30 / 1,000 repositories

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## GitHub Repo Scraper – Trending, Search, Releases & Issues

Get **GitHub repository data as clean JSON** from the **official GitHub REST API**: trending new repositories sorted by stars, repository search with filters, all repos of a user or organization, or details for your own list of repos. Each row has stars, forks, topics, license, language, dates and stars per day. You can add a language breakdown, the contributors count, **releases** (with asset download counts) and **issues / pull requests**.

It **works without a GitHub token**: the default run costs a single API request. Add your own token for large jobs. Price: **$0.30 per 1,000 repositories**.

### What it does

- **Trending repos.** Repositories created in the last N days, sorted by stars, optionally by language or topic. It is built on the official search API (`created:>=…`, `sort=stars`), so it does not break when the github.com/trending page changes.
- **Repository search.** Keywords plus any GitHub qualifier (`org:`, `license:`, `in:name`, `good-first-issues:>5`…), and filters for language, topic, star range, created and pushed dates, sort and order. Up to 1,000 results per query (GitHub's limit).
- **Owner mode.** All public repos of users or organizations (`apify`, `microsoft`, profile URLs), most stars first or by last push, creation or name. Forks and archived repos can be included or left out.
- **Repos mode.** Details for a list of `owner/name` entries or GitHub URLs, including real watchers and the fork parent.
- **Extras per repo (optional):**
  - `languages` (bytes) and `languagePercent`
  - `contributorsCount`
  - releases as rows: tag, date, pre-release, assets and download counts, plus `latestRelease` on the repo row
  - issues and/or PRs as rows: state, labels, `updated since`, author login, comments, reactions, merged date
- **Rate-limit aware.** Uses `per_page=100`, makes no needless calls and reports the remaining quota in every run. It waits for the per-minute search limit to reset. When the hourly limit runs out, it stops cleanly, keeps what was already saved and tells you when the limit resets.
- **Conditional requests (ETag)** for scheduled runs with a token: unchanged data comes back as HTTP 304, which does not count against your limit.
- **Privacy by design.** The Actor collects and outputs **no email addresses**, as GitHub's terms require. Email addresses inside issue or release text are replaced with `[email removed]`. Only public logins appear.

### Use cases

- **Trend scouting:** a daily or weekly "new and rising" list per language for newsletters, VC deal flow, devrel or content ideas (`starsPerDay` helps).
- **Open-source due diligence:** license, activity (`pushedAt`), contributors, release cadence and open issues for a list of dependencies.
- **Competitor and ecosystem tracking:** every repo of an organization, with stars and releases over time (schedule it and compare datasets).
- **Release monitoring:** the newest releases of the tools you depend on, with asset download counts.
- **Issue mining:** `good first issue` lists, bug reports since a date, or PR throughput for a project.
- **AI agents and RAG:** structured repo metadata for agents that pick libraries or answer "what is popular for X".

### Input example

```json
{
  "mode": "trending",
  "trendingDays": 7,
  "language": "Python",
  "maxResults": 50,
  "includeReleases": true,
  "maxReleasesPerRepo": 3
}
```

The default input (trending, last 7 days, 25 repos) needs one API request and takes about a second.

| Field | Default | Notes |
|---|---|---|
| `mode` | `trending` | `trending`, `search`, `owner`, `repos` |
| `trendingDays` | 7 | Trending: created in the last N days |
| `searchQuery`, `language`, `topic` | – | Search filters (also used by Trending) |
| `minStars` / `maxStars` | – | Star range |
| `createdAfter` / `createdBefore` / `pushedAfter` | – | `2026-09-01` or `30 days` |
| `sort` / `order` | stars / desc | Search mode |
| `maxResults` | 25 | Trending / Search, max 1,000 |
| `owners`, `ownerSort`, `maxReposPerOwner` | –, stars, 100 | Owner mode |
| `repositories` | – | Repos mode: `owner/name` or URLs |
| `includeForks` / `excludeArchived` | false / false | |
| `includeLanguages` / `includeContributorsCount` / `fetchFullDetails` | false | +1 API request per repo each |
| `includeReleases`, `maxReleasesPerRepo` | false, 10 | Release rows |
| `includeIssues`, `issueType`, `issueState`, `issueLabels`, `issuesSince`, `issueSort`, `maxIssuesPerRepo` | false, both, open, –, –, created, 30 | Issue / PR rows |
| `includeBodies`, `maxBodyChars` | false, 3000 | Release notes / issue text (emails removed) |
| `githubToken` | – | Optional, stored encrypted |
| `maxWaitForRateLimitSecs` | 90 | Wait if the limit resets within this time |
| `useConditionalRequests` | false | ETag cache for scheduled token runs |

### Output example

One dataset row per repository (`type: repo`), plus optional `release` and `issue` rows. The Output tab has **Repositories**, **Repo details & languages**, **Releases**, **Issues & pull requests** and **Errors** views.

```json
{
  "type": "repo",
  "fullName": "apify/crawlee",
  "owner": "apify",
  "ownerType": "Organization",
  "url": "https://github.com/apify/crawlee",
  "description": "Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. …",
  "homepage": "https://crawlee.dev",
  "primaryLanguage": "TypeScript",
  "topics": ["apify", "crawler", "playwright", "puppeteer", "scraping", "typescript", "web-scraping"],
  "license": "Apache-2.0",
  "stars": 25976,
  "forks": 1684,
  "watchers": 136,
  "openIssuesAndPrs": 134,
  "defaultBranch": "master",
  "isFork": false,
  "isArchived": false,
  "createdAt": "2016-08-26T18:35:03Z",
  "pushedAt": "2026-10-02T13:38:54Z",
  "ageDays": 3689.6,
  "starsPerDay": 7.04,
  "languagePercent": {"TypeScript": 62.1, "MDX": 31.0, "JavaScript": 5.4, "CSS": 0.9, "Dockerfile": 0.7},
  "contributorsCount": 139,
  "latestRelease": {"tagName": "v3.18.2", "publishedAt": "2026-09-29T09:58:35Z", "url": "https://github.com/apify/crawlee/releases/tag/v3.18.2"},
  "notes": null,
  "fetchedAt": "2026-10-03T09:37:20Z"
}
```

- **Release rows:** `repoFullName`, `tagName`, `name`, `publishedAt`, `isPrerelease`, `authorLogin`, `assets[]` (`name`, `sizeBytes`, `downloadCount`, `url`), `totalDownloads`, optional `body`.
- **Issue rows:** `repoFullName`, `number`, `isPullRequest`, `title`, `state`, `stateReason`, `labels`, `assignees`, `authorLogin`, `authorAssociation`, `comments`, `reactions`, `mergedAt`, `createdAt`, `updatedAt`, `closedAt`, optional `body`.
- **Errors:** `type: error` rows with `input`, `httpStatus`, `error` (repo not found or private, unknown user, invalid input). Not charged.
- **`notes`** on a repo row explains anything missing, e.g. "languages skipped: GitHub core rate limit reached".
- The key-value store record **`OUTPUT`** holds the run summary: counts, total search matches, API requests used, **remaining rate limit and reset time**, and whether the run stopped at the rate limit or at your max charge.

See `SAMPLE_OUTPUT.json` for full rows.

### Pricing

Pay per event. You pay only for rows saved.

| Event | Price |
|---|---|
| Repository | **$0.0003** ($0.30 / 1,000) |
| Release (optional) | $0.0001 ($0.10 / 1,000) |
| Issue or pull request (optional) | $0.0001 ($0.10 / 1,000) |
| Actor start | $0.00005 |

Languages, contributors count and full details are included in the repository price.

Compared with other Store Actors (public Store prices, 2026-10-01):

| Actor | Price per 1,000 repos |
|---|---|
| **GitHub Repo Scraper (this Actor)** | **$0.30** |
| dami_studio | $0.90 |
| automation-lab/github-trending-scraper | $1.15 |
| apivault | $1.90 |

Set **Maximum cost per run** in the run options to cap spending. The Actor checks the budget before each batch, makes no API calls for rows it cannot save, and stops cleanly.

### FAQ

**Do I need a GitHub token?**
No, not for the default run and small jobs. Without a token GitHub allows 60 requests per hour and 10 searches per minute per IP, and on Apify that IP is shared with other users. Search and Trending need only 1 request per 100 repos. Extras (languages, contributors, releases, issues) need about 1 request per repo each. For those, create a token at github.com → Settings → Developer settings → Personal access tokens (a classic token with no scopes, or a fine-grained token with public read access) and paste it into `githubToken`. That gives you 5,000 requests per hour.

**What happens when the rate limit is reached?**
The search limit resets every minute, so the Actor waits (up to `maxWaitForRateLimitSecs`). If the hourly limit is used up, the Actor stops cleanly. Repos already found by search are still saved, with a note on the skipped extras. The status message and `OUTPUT` show the reset time. You are never charged for data that was not saved.

**Is this the same as github.com/trending?**
It is a close, stable alternative: the most-starred repositories *created* in the last N days. GitHub's trending page uses a private ranking of recent star activity and has no API. For "rising" projects, sort by `starsPerDay`.

**Why at most 1,000 results per search?**
GitHub's search API returns at most 1,000 results per query. Split big jobs by star ranges (`minStars`/`maxStars`) or date ranges.

**Does it collect emails?**
No. GitHub's Acceptable Use Policies forbid using its API for spam or for selling users' personal information, so no email address is read or output. Commit author emails are never requested, and email addresses inside issue or release text are removed.

**Private repositories?**
Repos mode returns private repos your token can read. Without access you get a "not found (or private)" error row, free of charge.

**Is it legal?**
It uses only the official GitHub REST API, within GitHub's documented rate limits, with an identifying User-Agent. Follow GitHub's Terms of Service and each repository's license when you reuse the data.

### Use with AI agents / Apify MCP

- **Apify MCP server:** add `gazidev/github-repo-data` to your MCP client (Claude Desktop, Cursor, VS Code) via `https://mcp.apify.com?actors=gazidev/github-repo-data`. An agent can call it with `{"mode":"search","searchQuery":"vector database","language":"Rust","maxResults":10}` and get comparable repo stats as JSON.
- **API:** `POST https://api.apify.com/v2/acts/gazidev~github-repo-data/run-sync-get-dataset-items?token=...` with the input JSON returns the rows directly. It works well as a "GitHub tool" for LangChain or LlamaIndex agents.
- **Scheduled reports:** create a Task (e.g. trending Python, last 1 day), schedule it daily, and connect Slack, Google Sheets or a webhook.

### More developer and data tools

- [Hacker News Data](https://apify.com/gazidev/hacker-news-data): stories, comments and Who-is-hiring from the official HN APIs.
- [RSS Feed Reader](https://apify.com/gazidev/rss-feed-reader): any RSS, Atom or JSON Feed (GitHub release feeds too) to JSON, with only-new mode.
- [Tech Stack Detector](https://apify.com/gazidev/tech-stack-detector): which technologies a website uses.
- [Website to Markdown](https://apify.com/gazidev/website-to-markdown): docs sites to clean Markdown for RAG.

# Actor input Schema

## `mode` (type: `string`):

**Trending**: repositories created in the last N days, sorted by stars (a 'GitHub trending' style list built with the official search API). **Search**: any GitHub repository search with filters. **Owner**: all public repositories of users or organizations. **Repos**: details for a list of repositories (owner/name or URLs).

## `trendingDays` (type: `integer`):

Trending mode only. 1 = today's new repos, 7 = this week, 30 = this month.

## `searchQuery` (type: `string`):

Trending and Search modes. Free text (`web scraping`) and any GitHub qualifier (`in:name`, `org:apify`, `license:mit`, `good-first-issues:>5`, `is:public`). Combined with the filters below.

## `language` (type: `string`):

Primary language, e.g. `Python`, `TypeScript`, `Rust`. Several, comma-separated, means all must match, so use one.

## `topic` (type: `string`):

Repository topic, e.g. `llm`, `mcp`, `web-scraping`.

## `minStars` (type: `integer`):

0 = no minimum.

## `maxStars` (type: `integer`):

0 = no maximum. Useful to find small, fast-growing projects.

## `createdAfter` (type: `string`):

Search mode. Absolute (`2026-01-01`) or relative (`30 days`). Trending mode uses 'last N days' instead.

## `createdBefore` (type: `string`):

Search mode. Absolute or relative date.

## `pushedAfter` (type: `string`):

Only actively maintained repos: last commit pushed on or after this date (`2026-09-01` or `90 days`).

## `sort` (type: `string`):

Trending mode always sorts by stars.

## `order` (type: `string`):

Descending = most stars / newest first.

## `maxResults` (type: `integer`):

GitHub returns at most 1,000 results per search. Split big jobs by date or star ranges.

## `owners` (type: `array`):

Logins or profile URLs, e.g. `apify`, `microsoft`, `https://github.com/torvalds`.

## `ownerSort` (type: `string`):

Most stars first uses the search API (1 request per 100 repos, max 1,000 per owner). The others list repos directly.

## `maxReposPerOwner` (type: `integer`):

Owner mode only.

## `repositories` (type: `array`):

`owner/name` or GitHub URLs, e.g. `apify/crawlee`, `https://github.com/psf/requests`. One API request per repo (plus the extras below).

## `includeForks` (type: `boolean`):

Search and Owner modes leave forks out by default.

## `excludeArchived` (type: `boolean`):

Leave out archived (read-only) repositories in Search and Owner modes.

## `includeLanguages` (type: `boolean`):

Adds `languages` (bytes per language) and `languagePercent`. +1 API request per repo.

## `includeContributorsCount` (type: `boolean`):

Adds `contributorsCount` (contributors with a GitHub account). +1 API request per repo. GitHub does not count very large repos (e.g. linux).

## `fetchFullDetails` (type: `boolean`):

Trending / Search / Owner results already include stars, forks, topics, license and dates. Turn on for real `watchers` and the fork `parentFullName`. +1 API request per repo. Repos mode always fetches details.

## `includeReleases` (type: `boolean`):

Saves the newest releases of each repo as separate rows (`type: release`): tag, date, assets and download counts. Adds `latestRelease` to the repo row. Charged per release row.

## `maxReleasesPerRepo` (type: `integer`):

Newest first. One API request per 100 releases.

## `includeIssues` (type: `boolean`):

Saves issues and/or pull requests as separate rows (`type: issue`). Charged per row.

## `issueType` (type: `string`):

GitHub's issues endpoint returns both; picking one filters the list.

## `issueState` (type: `string`):

Open, closed or all.

## `issueLabels` (type: `string`):

Comma-separated; all must match, e.g. `bug` or `good first issue`.

## `issuesSince` (type: `string`):

Only issues/PRs updated on or after this date (`2026-09-01` or `7 days`).

## `issueSort` (type: `string`):

Newest first by this field.

## `maxIssuesPerRepo` (type: `integer`):

Newest first (see 'Issue order'). One API request per 100 rows.

## `includeBodies` (type: `boolean`):

Adds `body` (Markdown) to release and issue rows. Email addresses in the text are removed.

## `maxBodyChars` (type: `integer`):

Longer texts are cut. 0 = no limit.

## `githubToken` (type: `string`):

Optional personal access token (classic with no scopes, or fine-grained with public read access). Without a token GitHub allows 60 requests/hour and 10 searches/minute per IP, which is enough for the default run and small jobs. With a token: 5,000 requests/hour and 30 searches/minute. Stored encrypted; used only for api.github.com.

## `maxWaitForRateLimitSecs` (type: `integer`):

If the limit resets within this time (the search limit resets every minute), the Actor waits and continues. Otherwise it stops cleanly, keeps everything already saved and reports the reset time.

## `useConditionalRequests` (type: `boolean`):

For scheduled runs with a token: remembers ETags in a named key-value store and sends If-None-Match. Unchanged data comes back as HTTP 304, which GitHub does not count against the token's limit. Adds key-value store operations, so leave off for one-off runs.

## `maxConcurrency` (type: `integer`):

GitHub asks API clients to avoid many concurrent requests. 4 is a safe default.

## Actor input object example

```json
{
  "mode": "trending",
  "trendingDays": 7,
  "sort": "stars",
  "order": "desc",
  "maxResults": 25,
  "owners": [
    "apify"
  ],
  "ownerSort": "stars",
  "maxReposPerOwner": 100,
  "repositories": [
    "apify/crawlee",
    "psf/requests"
  ],
  "includeForks": false,
  "excludeArchived": false,
  "includeLanguages": false,
  "includeContributorsCount": false,
  "fetchFullDetails": false,
  "includeReleases": false,
  "maxReleasesPerRepo": 10,
  "includeIssues": false,
  "issueType": "both",
  "issueState": "open",
  "issueSort": "created",
  "maxIssuesPerRepo": 30,
  "includeBodies": false,
  "maxBodyChars": 3000,
  "maxWaitForRateLimitSecs": 90,
  "useConditionalRequests": false,
  "maxConcurrency": 4
}
```

# Actor output Schema

## `repos` (type: `string`):

No description

## `details` (type: `string`):

No description

## `releases` (type: `string`):

No description

## `issues` (type: `string`):

No description

## `errors` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "trending",
    "trendingDays": 7,
    "maxResults": 25,
    "owners": [
        "apify"
    ],
    "repositories": [
        "apify/crawlee",
        "psf/requests"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("gazidev/github-repo-data").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "trending",
    "trendingDays": 7,
    "maxResults": 25,
    "owners": ["apify"],
    "repositories": [
        "apify/crawlee",
        "psf/requests",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("gazidev/github-repo-data").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "trending",
  "trendingDays": 7,
  "maxResults": 25,
  "owners": [
    "apify"
  ],
  "repositories": [
    "apify/crawlee",
    "psf/requests"
  ]
}' |
apify call gazidev/github-repo-data --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gazidev/github-repo-data"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZjEtFjtf84n97rbnJ/builds/XmY5ZLIibGfNcl72g/openapi.json
