GitHub Repo Scraper – Trending, Search, Releases & Issues API avatar

GitHub Repo Scraper – Trending, Search, Releases & Issues API

Pricing

from $0.30 / 1,000 repositories

Go to Apify Store
GitHub Repo Scraper – Trending, Search, Releases & Issues API

GitHub Repo Scraper – Trending, Search, Releases & Issues API

Get GitHub repository data from the official REST API: trending new repos by stars, search with filters, all repos of a user or org, or details for a list. Stars, forks, topics, license, languages, contributors count, releases and issues/PRs. Works without a token; no emails collected.

Pricing

from $0.30 / 1,000 repositories

Rating

0.0

(0)

Developer

Cemal Atakli

Cemal Atakli

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

GitHub Repo Scraper – Trending, Search, Releases & Issues

Get GitHub repository data as clean JSON from the official GitHub REST API: trending new repositories sorted by stars, repository search with filters, all repos of a user or organization, or details for your own list of repos. Each row has stars, forks, topics, license, language, dates and stars per day. You can add a language breakdown, the contributors count, releases (with asset download counts) and issues / pull requests.

It works without a GitHub token: the default run costs a single API request. Add your own token for large jobs. Price: $0.30 per 1,000 repositories.

What it does

  • Trending repos. Repositories created in the last N days, sorted by stars, optionally by language or topic. It is built on the official search API (created:>=…, sort=stars), so it does not break when the github.com/trending page changes.
  • Repository search. Keywords plus any GitHub qualifier (org:, license:, in:name, good-first-issues:>5…), and filters for language, topic, star range, created and pushed dates, sort and order. Up to 1,000 results per query (GitHub's limit).
  • Owner mode. All public repos of users or organizations (apify, microsoft, profile URLs), most stars first or by last push, creation or name. Forks and archived repos can be included or left out.
  • Repos mode. Details for a list of owner/name entries or GitHub URLs, including real watchers and the fork parent.
  • Extras per repo (optional):
    • languages (bytes) and languagePercent
    • contributorsCount
    • releases as rows: tag, date, pre-release, assets and download counts, plus latestRelease on the repo row
    • issues and/or PRs as rows: state, labels, updated since, author login, comments, reactions, merged date
  • Rate-limit aware. Uses per_page=100, makes no needless calls and reports the remaining quota in every run. It waits for the per-minute search limit to reset. When the hourly limit runs out, it stops cleanly, keeps what was already saved and tells you when the limit resets.
  • Conditional requests (ETag) for scheduled runs with a token: unchanged data comes back as HTTP 304, which does not count against your limit.
  • Privacy by design. The Actor collects and outputs no email addresses, as GitHub's terms require. Email addresses inside issue or release text are replaced with [email removed]. Only public logins appear.

Use cases

  • Trend scouting: a daily or weekly "new and rising" list per language for newsletters, VC deal flow, devrel or content ideas (starsPerDay helps).
  • Open-source due diligence: license, activity (pushedAt), contributors, release cadence and open issues for a list of dependencies.
  • Competitor and ecosystem tracking: every repo of an organization, with stars and releases over time (schedule it and compare datasets).
  • Release monitoring: the newest releases of the tools you depend on, with asset download counts.
  • Issue mining: good first issue lists, bug reports since a date, or PR throughput for a project.
  • AI agents and RAG: structured repo metadata for agents that pick libraries or answer "what is popular for X".

Input example

{
"mode": "trending",
"trendingDays": 7,
"language": "Python",
"maxResults": 50,
"includeReleases": true,
"maxReleasesPerRepo": 3
}

The default input (trending, last 7 days, 25 repos) needs one API request and takes about a second.

FieldDefaultNotes
modetrendingtrending, search, owner, repos
trendingDays7Trending: created in the last N days
searchQuery, language, topic–Search filters (also used by Trending)
minStars / maxStars–Star range
createdAfter / createdBefore / pushedAfter–2026-09-01 or 30 days
sort / orderstars / descSearch mode
maxResults25Trending / Search, max 1,000
owners, ownerSort, maxReposPerOwner–, stars, 100Owner mode
repositories–Repos mode: owner/name or URLs
includeForks / excludeArchivedfalse / false
includeLanguages / includeContributorsCount / fetchFullDetailsfalse+1 API request per repo each
includeReleases, maxReleasesPerRepofalse, 10Release rows
includeIssues, issueType, issueState, issueLabels, issuesSince, issueSort, maxIssuesPerRepofalse, both, open, –, –, created, 30Issue / PR rows
includeBodies, maxBodyCharsfalse, 3000Release notes / issue text (emails removed)
githubToken–Optional, stored encrypted
maxWaitForRateLimitSecs90Wait if the limit resets within this time
useConditionalRequestsfalseETag cache for scheduled token runs

Output example

One dataset row per repository (type: repo), plus optional release and issue rows. The Output tab has Repositories, Repo details & languages, Releases, Issues & pull requests and Errors views.

{
"type": "repo",
"fullName": "apify/crawlee",
"owner": "apify",
"ownerType": "Organization",
"url": "https://github.com/apify/crawlee",
"description": "Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. …",
"homepage": "https://crawlee.dev",
"primaryLanguage": "TypeScript",
"topics": ["apify", "crawler", "playwright", "puppeteer", "scraping", "typescript", "web-scraping"],
"license": "Apache-2.0",
"stars": 25976,
"forks": 1684,
"watchers": 136,
"openIssuesAndPrs": 134,
"defaultBranch": "master",
"isFork": false,
"isArchived": false,
"createdAt": "2016-08-26T18:35:03Z",
"pushedAt": "2026-10-02T13:38:54Z",
"ageDays": 3689.6,
"starsPerDay": 7.04,
"languagePercent": {"TypeScript": 62.1, "MDX": 31.0, "JavaScript": 5.4, "CSS": 0.9, "Dockerfile": 0.7},
"contributorsCount": 139,
"latestRelease": {"tagName": "v3.18.2", "publishedAt": "2026-09-29T09:58:35Z", "url": "https://github.com/apify/crawlee/releases/tag/v3.18.2"},
"notes": null,
"fetchedAt": "2026-10-03T09:37:20Z"
}
  • Release rows: repoFullName, tagName, name, publishedAt, isPrerelease, authorLogin, assets[] (name, sizeBytes, downloadCount, url), totalDownloads, optional body.
  • Issue rows: repoFullName, number, isPullRequest, title, state, stateReason, labels, assignees, authorLogin, authorAssociation, comments, reactions, mergedAt, createdAt, updatedAt, closedAt, optional body.
  • Errors: type: error rows with input, httpStatus, error (repo not found or private, unknown user, invalid input). Not charged.
  • notes on a repo row explains anything missing, e.g. "languages skipped: GitHub core rate limit reached".
  • The key-value store record OUTPUT holds the run summary: counts, total search matches, API requests used, remaining rate limit and reset time, and whether the run stopped at the rate limit or at your max charge.

See SAMPLE_OUTPUT.json for full rows.

Pricing

Pay per event. You pay only for rows saved.

EventPrice
Repository$0.0003 ($0.30 / 1,000)
Release (optional)$0.0001 ($0.10 / 1,000)
Issue or pull request (optional)$0.0001 ($0.10 / 1,000)
Actor start$0.00005

Languages, contributors count and full details are included in the repository price.

Compared with other Store Actors (public Store prices, 2026-10-01):

ActorPrice per 1,000 repos
GitHub Repo Scraper (this Actor)$0.30
dami_studio$0.90
automation-lab/github-trending-scraper$1.15
apivault$1.90

Set Maximum cost per run in the run options to cap spending. The Actor checks the budget before each batch, makes no API calls for rows it cannot save, and stops cleanly.

FAQ

Do I need a GitHub token? No, not for the default run and small jobs. Without a token GitHub allows 60 requests per hour and 10 searches per minute per IP, and on Apify that IP is shared with other users. Search and Trending need only 1 request per 100 repos. Extras (languages, contributors, releases, issues) need about 1 request per repo each. For those, create a token at github.com → Settings → Developer settings → Personal access tokens (a classic token with no scopes, or a fine-grained token with public read access) and paste it into githubToken. That gives you 5,000 requests per hour.

What happens when the rate limit is reached? The search limit resets every minute, so the Actor waits (up to maxWaitForRateLimitSecs). If the hourly limit is used up, the Actor stops cleanly. Repos already found by search are still saved, with a note on the skipped extras. The status message and OUTPUT show the reset time. You are never charged for data that was not saved.

Is this the same as github.com/trending? It is a close, stable alternative: the most-starred repositories created in the last N days. GitHub's trending page uses a private ranking of recent star activity and has no API. For "rising" projects, sort by starsPerDay.

Why at most 1,000 results per search? GitHub's search API returns at most 1,000 results per query. Split big jobs by star ranges (minStars/maxStars) or date ranges.

Does it collect emails? No. GitHub's Acceptable Use Policies forbid using its API for spam or for selling users' personal information, so no email address is read or output. Commit author emails are never requested, and email addresses inside issue or release text are removed.

Private repositories? Repos mode returns private repos your token can read. Without access you get a "not found (or private)" error row, free of charge.

Is it legal? It uses only the official GitHub REST API, within GitHub's documented rate limits, with an identifying User-Agent. Follow GitHub's Terms of Service and each repository's license when you reuse the data.

Use with AI agents / Apify MCP

  • Apify MCP server: add gazidev/github-repo-data to your MCP client (Claude Desktop, Cursor, VS Code) via https://mcp.apify.com?actors=gazidev/github-repo-data. An agent can call it with {"mode":"search","searchQuery":"vector database","language":"Rust","maxResults":10} and get comparable repo stats as JSON.
  • API: POST https://api.apify.com/v2/acts/gazidev~github-repo-data/run-sync-get-dataset-items?token=... with the input JSON returns the rows directly. It works well as a "GitHub tool" for LangChain or LlamaIndex agents.
  • Scheduled reports: create a Task (e.g. trending Python, last 1 day), schedule it daily, and connect Slack, Google Sheets or a webhook.

More developer and data tools