GitHub Repo Scraper – Trending, Search, Releases & Issues API
Pricing
from $0.30 / 1,000 repositories
GitHub Repo Scraper – Trending, Search, Releases & Issues API
Get GitHub repository data from the official REST API: trending new repos by stars, search with filters, all repos of a user or org, or details for a list. Stars, forks, topics, license, languages, contributors count, releases and issues/PRs. Works without a token; no emails collected.
Pricing
from $0.30 / 1,000 repositories
Rating
0.0
(0)
Developer
Cemal Atakli
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
GitHub Repo Scraper – Trending, Search, Releases & Issues
Get GitHub repository data as clean JSON from the official GitHub REST API: trending new repositories sorted by stars, repository search with filters, all repos of a user or organization, or details for your own list of repos. Each row has stars, forks, topics, license, language, dates and stars per day. You can add a language breakdown, the contributors count, releases (with asset download counts) and issues / pull requests.
It works without a GitHub token: the default run costs a single API request. Add your own token for large jobs. Price: $0.30 per 1,000 repositories.
What it does
- Trending repos. Repositories created in the last N days, sorted by stars, optionally by language or topic. It is built on the official search API (
created:>=…,sort=stars), so it does not break when the github.com/trending page changes. - Repository search. Keywords plus any GitHub qualifier (
org:,license:,in:name,good-first-issues:>5…), and filters for language, topic, star range, created and pushed dates, sort and order. Up to 1,000 results per query (GitHub's limit). - Owner mode. All public repos of users or organizations (
apify,microsoft, profile URLs), most stars first or by last push, creation or name. Forks and archived repos can be included or left out. - Repos mode. Details for a list of
owner/nameentries or GitHub URLs, including real watchers and the fork parent. - Extras per repo (optional):
languages(bytes) andlanguagePercentcontributorsCount- releases as rows: tag, date, pre-release, assets and download counts, plus
latestReleaseon the repo row - issues and/or PRs as rows: state, labels,
updated since, author login, comments, reactions, merged date
- Rate-limit aware. Uses
per_page=100, makes no needless calls and reports the remaining quota in every run. It waits for the per-minute search limit to reset. When the hourly limit runs out, it stops cleanly, keeps what was already saved and tells you when the limit resets. - Conditional requests (ETag) for scheduled runs with a token: unchanged data comes back as HTTP 304, which does not count against your limit.
- Privacy by design. The Actor collects and outputs no email addresses, as GitHub's terms require. Email addresses inside issue or release text are replaced with
[email removed]. Only public logins appear.
Use cases
- Trend scouting: a daily or weekly "new and rising" list per language for newsletters, VC deal flow, devrel or content ideas (
starsPerDayhelps). - Open-source due diligence: license, activity (
pushedAt), contributors, release cadence and open issues for a list of dependencies. - Competitor and ecosystem tracking: every repo of an organization, with stars and releases over time (schedule it and compare datasets).
- Release monitoring: the newest releases of the tools you depend on, with asset download counts.
- Issue mining:
good first issuelists, bug reports since a date, or PR throughput for a project. - AI agents and RAG: structured repo metadata for agents that pick libraries or answer "what is popular for X".
Input example
{"mode": "trending","trendingDays": 7,"language": "Python","maxResults": 50,"includeReleases": true,"maxReleasesPerRepo": 3}
The default input (trending, last 7 days, 25 repos) needs one API request and takes about a second.
| Field | Default | Notes |
|---|---|---|
mode | trending | trending, search, owner, repos |
trendingDays | 7 | Trending: created in the last N days |
searchQuery, language, topic | – | Search filters (also used by Trending) |
minStars / maxStars | – | Star range |
createdAfter / createdBefore / pushedAfter | – | 2026-09-01 or 30 days |
sort / order | stars / desc | Search mode |
maxResults | 25 | Trending / Search, max 1,000 |
owners, ownerSort, maxReposPerOwner | –, stars, 100 | Owner mode |
repositories | – | Repos mode: owner/name or URLs |
includeForks / excludeArchived | false / false | |
includeLanguages / includeContributorsCount / fetchFullDetails | false | +1 API request per repo each |
includeReleases, maxReleasesPerRepo | false, 10 | Release rows |
includeIssues, issueType, issueState, issueLabels, issuesSince, issueSort, maxIssuesPerRepo | false, both, open, –, –, created, 30 | Issue / PR rows |
includeBodies, maxBodyChars | false, 3000 | Release notes / issue text (emails removed) |
githubToken | – | Optional, stored encrypted |
maxWaitForRateLimitSecs | 90 | Wait if the limit resets within this time |
useConditionalRequests | false | ETag cache for scheduled token runs |
Output example
One dataset row per repository (type: repo), plus optional release and issue rows. The Output tab has Repositories, Repo details & languages, Releases, Issues & pull requests and Errors views.
{"type": "repo","fullName": "apify/crawlee","owner": "apify","ownerType": "Organization","url": "https://github.com/apify/crawlee","description": "Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. …","homepage": "https://crawlee.dev","primaryLanguage": "TypeScript","topics": ["apify", "crawler", "playwright", "puppeteer", "scraping", "typescript", "web-scraping"],"license": "Apache-2.0","stars": 25976,"forks": 1684,"watchers": 136,"openIssuesAndPrs": 134,"defaultBranch": "master","isFork": false,"isArchived": false,"createdAt": "2016-08-26T18:35:03Z","pushedAt": "2026-10-02T13:38:54Z","ageDays": 3689.6,"starsPerDay": 7.04,"languagePercent": {"TypeScript": 62.1, "MDX": 31.0, "JavaScript": 5.4, "CSS": 0.9, "Dockerfile": 0.7},"contributorsCount": 139,"latestRelease": {"tagName": "v3.18.2", "publishedAt": "2026-09-29T09:58:35Z", "url": "https://github.com/apify/crawlee/releases/tag/v3.18.2"},"notes": null,"fetchedAt": "2026-10-03T09:37:20Z"}
- Release rows:
repoFullName,tagName,name,publishedAt,isPrerelease,authorLogin,assets[](name,sizeBytes,downloadCount,url),totalDownloads, optionalbody. - Issue rows:
repoFullName,number,isPullRequest,title,state,stateReason,labels,assignees,authorLogin,authorAssociation,comments,reactions,mergedAt,createdAt,updatedAt,closedAt, optionalbody. - Errors:
type: errorrows withinput,httpStatus,error(repo not found or private, unknown user, invalid input). Not charged. noteson a repo row explains anything missing, e.g. "languages skipped: GitHub core rate limit reached".- The key-value store record
OUTPUTholds the run summary: counts, total search matches, API requests used, remaining rate limit and reset time, and whether the run stopped at the rate limit or at your max charge.
See SAMPLE_OUTPUT.json for full rows.
Pricing
Pay per event. You pay only for rows saved.
| Event | Price |
|---|---|
| Repository | $0.0003 ($0.30 / 1,000) |
| Release (optional) | $0.0001 ($0.10 / 1,000) |
| Issue or pull request (optional) | $0.0001 ($0.10 / 1,000) |
| Actor start | $0.00005 |
Languages, contributors count and full details are included in the repository price.
Compared with other Store Actors (public Store prices, 2026-10-01):
| Actor | Price per 1,000 repos |
|---|---|
| GitHub Repo Scraper (this Actor) | $0.30 |
| dami_studio | $0.90 |
| automation-lab/github-trending-scraper | $1.15 |
| apivault | $1.90 |
Set Maximum cost per run in the run options to cap spending. The Actor checks the budget before each batch, makes no API calls for rows it cannot save, and stops cleanly.
FAQ
Do I need a GitHub token?
No, not for the default run and small jobs. Without a token GitHub allows 60 requests per hour and 10 searches per minute per IP, and on Apify that IP is shared with other users. Search and Trending need only 1 request per 100 repos. Extras (languages, contributors, releases, issues) need about 1 request per repo each. For those, create a token at github.com → Settings → Developer settings → Personal access tokens (a classic token with no scopes, or a fine-grained token with public read access) and paste it into githubToken. That gives you 5,000 requests per hour.
What happens when the rate limit is reached?
The search limit resets every minute, so the Actor waits (up to maxWaitForRateLimitSecs). If the hourly limit is used up, the Actor stops cleanly. Repos already found by search are still saved, with a note on the skipped extras. The status message and OUTPUT show the reset time. You are never charged for data that was not saved.
Is this the same as github.com/trending?
It is a close, stable alternative: the most-starred repositories created in the last N days. GitHub's trending page uses a private ranking of recent star activity and has no API. For "rising" projects, sort by starsPerDay.
Why at most 1,000 results per search?
GitHub's search API returns at most 1,000 results per query. Split big jobs by star ranges (minStars/maxStars) or date ranges.
Does it collect emails? No. GitHub's Acceptable Use Policies forbid using its API for spam or for selling users' personal information, so no email address is read or output. Commit author emails are never requested, and email addresses inside issue or release text are removed.
Private repositories? Repos mode returns private repos your token can read. Without access you get a "not found (or private)" error row, free of charge.
Is it legal? It uses only the official GitHub REST API, within GitHub's documented rate limits, with an identifying User-Agent. Follow GitHub's Terms of Service and each repository's license when you reuse the data.
Use with AI agents / Apify MCP
- Apify MCP server: add
gazidev/github-repo-datato your MCP client (Claude Desktop, Cursor, VS Code) viahttps://mcp.apify.com?actors=gazidev/github-repo-data. An agent can call it with{"mode":"search","searchQuery":"vector database","language":"Rust","maxResults":10}and get comparable repo stats as JSON. - API:
POST https://api.apify.com/v2/acts/gazidev~github-repo-data/run-sync-get-dataset-items?token=...with the input JSON returns the rows directly. It works well as a "GitHub tool" for LangChain or LlamaIndex agents. - Scheduled reports: create a Task (e.g. trending Python, last 1 day), schedule it daily, and connect Slack, Google Sheets or a webhook.
More developer and data tools
- Hacker News Data: stories, comments and Who-is-hiring from the official HN APIs.
- RSS Feed Reader: any RSS, Atom or JSON Feed (GitHub release feeds too) to JSON, with only-new mode.
- Tech Stack Detector: which technologies a website uses.
- Website to Markdown: docs sites to clean Markdown for RAG.