GitHub Repo Scraper: Stars, Releases, Licenses, Dependencies
Pricing
$4.00 / 1,000 repositories
GitHub Repo Scraper: Stars, Releases, Licenses, Dependencies
Returns one row per GitHub repository with stars, forks, topics, license, languages with share, latest commit, latest release and release count, and package.json, pyproject.toml or go.mod details. Repository lists or GitHub search, optional token, onlyNew for release monitoring.
Pricing
$4.00 / 1,000 repositories
Rating
0.0
(0)
Developer
Viktor Wiberg
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
10 hours ago
Last modified
Categories
Share
A GitHub scraper that returns one row per repository with the facts you would otherwise collect by hand from several pages: GitHub stars, forks, topics, license, the languages and their share of the code, the latest commit on the default branch, the latest release and how many releases there are, and the package name, version and dependency count from package.json, pyproject.toml or go.mod.
Give it a list of repositories, a GitHub search such as topic:web-scraping stars:>500, or both. It reads the official GitHub REST API and works without a login. A GitHub token is optional and only raises the rate limit.
The actor returns repository data only. It does not read user profiles, email addresses, follower lists or the names of commit authors and contributors.
Example from a real run
Run MBr1GDZhdckCRZ8bB on Apify on 5 October 2026, without a token, with this input:
{"repos": ["cli/cli", "apify/crawlee-python", "microsoft/playwright", "psf/requests", "denoland/deno"],"includeReleases": true,"includeManifest": true}
It returned 5 rows in 5 seconds of run time and used 20 of the 60 GitHub requests allowed per hour without a token. The first row, with languages cut to two entries and source, sourceLicense and retrievedAt left out here to keep it short:
{"fullName": "cli/cli","owner": "cli","name": "cli","url": "https://github.com/cli/cli","status": "ok","error": null,"description": "GitHub’s official command line tool","homepage": "https://cli.github.com","topics": ["cli", "git", "github-api-v4", "golang"],"license": "MIT","licenseName": "MIT License","stars": 46544,"forks": 9121,"watchers": 1140,"openIssues": 1116,"defaultBranch": "trunk","primaryLanguage": "Go","languages": [{ "language": "Go", "bytes": 7892539, "share": 99.4 },{ "language": "Shell", "bytes": 41794, "share": 0.5 }],"isArchived": false,"isFork": false,"isTemplate": false,"sizeKb": 81873,"createdAt": "2019-10-03T15:24:53Z","updatedAt": "2026-10-05T02:41:52Z","pushedAt": "2026-10-02T23:13:52Z","lastCommitSha": "6fc1c29d5477bfe71da7af290eb481c0df7811f1","lastCommitDate": "2026-10-02T16:05:07Z","latestReleaseTag": "v2.102.0","latestReleaseName": "GitHub CLI 2.102.0","latestReleaseDate": "2026-09-30T02:40:02Z","latestReleasePrerelease": false,"latestReleaseUrl": "https://github.com/cli/cli/releases/tag/v2.102.0","releaseCount": 205,"manifests": [{"file": "go.mod","ecosystem": "go","packageName": "github.com/cli/cli/v2","version": null,"dependencies": 61,"devDependencies": 118}]}
The other four rows in the same run: apify/crawlee-python (Apache-2.0, release v1.10.3, pyproject.toml with 13 dependencies), microsoft/playwright (Apache-2.0, release v1.63.0, package.json), psf/requests (Apache-2.0, release v2.34.2, 19 releases) and denoland/deno (MIT, release v2.9.7, 390 releases, no manifest in the root since it is a Rust project).
A search run on the same day (run DgYbbzawDgrRmDqlb, input {"search": "topic:web-scraping stars:>500", "sort": "stars", "maxResults": 10}) found 133 matching repositories and returned the 10 with the most stars, starting with firecrawl/firecrawl, D4Vinci/Scrapling and unclecode/crawl4ai.
With empty input the actor reads apify/crawlee and apify/apify-sdk-js as an example (run yDFLkbdeyDRY2XEXW, 2 rows, 8 GitHub requests).
Input
| Field | Type | Default | Description |
|---|---|---|---|
repos | array | Repositories as owner/name or a github.com link, for example apify/crawlee or https://github.com/cli/cli. | |
search | string | A GitHub repository search with the same syntax as on github.com, for example topic:web-scraping stars:>500 or language:go stars:>10000 pushed:>2026-09-01. | |
sort | string | stars | Order of the search results: stars, forks, updated, help-wanted-issues or best-match. |
githubToken | string | Optional. Raises the limit from 60 to 5 000 GitHub requests per hour. See "GitHub token". | |
includeReleases | boolean | true | Latest release and number of releases. One extra request per repository. |
includeManifest | boolean | true | Package name, version and dependency counts from package.json, pyproject.toml and go.mod in the root of the default branch. |
maxResults | integer | 50 | Maximum number of repositories. Named repositories count first, then search results. 1 to 1 000. |
onlyNew | boolean | false | Return only repositories with a new push or a new release since the last run with the same input. See "Monitoring and scheduling". |
When both repos and search are set, the named repositories come first and the search fills up to maxResults. Without either, the two example repositories are read.
Output
| Field | Description |
|---|---|
fullName, owner, name, url | The repository and its link |
status, error | ok, or error with the reason (not found, private, rate limit reached). Error rows are not charged |
description, homepage, topics | As set by the maintainers |
license, licenseName | SPDX id and name as detected by GitHub, for example MIT. NOASSERTION means GitHub found a license file it could not classify. null means no license file |
stars, forks, watchers | Stargazers, forks and subscribers. watchers is null for rows from a search, since the search API does not return it |
openIssues | Open issues and pull requests together, as GitHub counts them |
defaultBranch, primaryLanguage | Default branch and the language GitHub shows for the repository |
languages | Every language with bytes of code and percent of the total, largest first |
isArchived, isFork, isTemplate | Repository flags |
sizeKb, createdAt, updatedAt, pushedAt | Size of the repository and its dates, UTC |
lastCommitSha, lastCommitDate | The newest commit on the default branch and its commit date. No author name or email |
latestReleaseTag, latestReleaseName, latestReleaseDate, latestReleasePrerelease, latestReleaseUrl | The most recently created release, which can be a prerelease (see the flag) |
releaseCount | Number of published releases. Tags without a release are not counted |
manifests | One entry per manifest found: file, ecosystem (npm, python, go), packageName, version, dependencies, devDependencies |
source, sourceLicense, retrievedAt | Attribution and the time of the run |
How dependencies are counted: in package.json, dependencies and devDependencies. In pyproject.toml, [project] dependencies, and as devDependencies the optional dependencies plus dependency groups (for Poetry projects, [tool.poetry.dependencies] without python, and the dev groups). In go.mod, direct requirements as dependencies and // indirect ones as devDependencies. A version written as dynamic in pyproject.toml comes back as null.
GitHub token
The actor works without a token. GitHub then allows 60 API requests per hour per IP address, and Apify's IP addresses are shared, so part of that hour may already be used by others when your run starts.
Requests per repository: 1 for the repository record (0 when it comes from a search), 1 for languages, 1 for the latest commit and 1 for releases when includeReleases is on. Manifests are read from raw.githubusercontent.com and do not count. So without a token:
- 50 named repositories need 200 requests, and a run reads about 14 of them before the limit.
- A search for 50 repositories needs 151 requests (1 search plus 3 per repository) and reads about 19.
- With
includeReleasesoff, each repository needs one request less.
When the limit is reached, the actor stops calmly: every repository it did not read gets a row with status error and the time the limit resets, the run finishes as succeeded and those rows are not charged. A short secondary limit (a Retry-After of up to 60 seconds) is waited out and retried.
With a token the limit is 5 000 requests per hour, enough for about 1 250 repositories. To create one that can only read public data:
- On GitHub, open Settings, Developer settings, Personal access tokens, Fine-grained tokens, and choose Generate new token.
- Give it a name and an expiry date. Under Repository access, choose Public repositories (read-only). Leave all permissions at No access.
- Generate it and paste it into
githubToken.
The field is marked secret, so Apify stores it encrypted and the actor never writes it to the log or the dataset. A token cannot read anything private here: the actor skips any repository that is not public. Use your own token only. GitHub's terms do not allow sharing tokens to get around its rate limits.
Monitoring and scheduling
Set onlyNew to true to watch a list of repositories on a schedule. The actor remembers what it has delivered for that input in a named key-value store in your Apify account (nightwave-state-github-repo-intelligence, one record per input). A repository is delivered again when its pushedAt time or its latest release tag changes, so you get a row after a new push or a new release, and nothing for repositories that have not changed. Only delivered rows are charged. The first run returns everything.
onlyNew, maxResults and githubToken are not part of the remembered input, so you can change them without losing the state. Changing any other field starts a fresh state. To start over, delete the record in the key-value store.
In a test on 5 October 2026, the input {"repos": ["apify/crawlee", "cli/cli", "psf/requests"], "onlyNew": true} returned 3 rows on the first run (run ZoikDDQWPorHclsPg) and 0 rows on a second run right after (run YXheiYG7rD2cuFqL6), since nothing had changed in between. A second run like that, with onlyNew and the same input, is not charged.
Example: every morning at 07:00 Swedish time, check the dependencies you care about for new releases. In Apify Console, open Schedules, create a schedule with the cron expression 0 7 * * * and add this actor with the input below.
{"repos": ["nodejs/node", "microsoft/playwright", "apify/crawlee", "psf/requests"],"onlyNew": true}
The same schedule through the Apify API:
curl -X POST "https://api.apify.com/v2/schedules?token=<YOUR_TOKEN>" \-H "Content-Type: application/json" \-d '{"name": "daily-release-check", "cronExpression": "0 7 * * *", "timezone": "Europe/Stockholm", "isEnabled": true, "isExclusive": true,"actions": [{"type": "RUN_ACTOR", "actorId": "nightwave-owner~github-repo-intelligence","runInput": {"contentType": "application/json; charset=utf-8", "body": "<the input above as a JSON string>"}}]}'
Add an Apify integration (email, Slack or a webhook) on the schedule to be told when a run has rows. A push to any branch counts as a change, so for busy repositories you will get a row most days. Compare latestReleaseTag with the previous row if you only care about releases.
Use cases
- Release monitoring of dependencies. Run the repositories your product depends on every day with
onlyNewand get a row only when one of them has pushed or published a new release. - Due diligence on open source. Before you adopt a project, check its maintenance activity (last commit, last release, release count, archived or not), its size and who stands behind it.
- License inventory. List the license of every repository in a dependency list or an organization, and find projects without a license or with one your policy does not allow.
- Market research for developer tools. Pull the top repositories for a topic or language with stars, forks, release cadence and package names.
Limitations
- GitHub's search returns at most 1 000 repositories per query. Split a large query, for example by
created:ranges, to cover more. releaseCountcounts releases, not git tags. Some projects only tag and never publish releases on GitHub; they havereleaseCount0.- The latest release is the most recently created one and can be a prerelease. Check
latestReleasePrerelease. - Manifests are read only from the root of the default branch. Monorepos often have a private root
package.json(for example@crawlee/rootwith no version), and projects with their manifest in a subfolder return an empty list. openIssuesincludes open pull requests, since that is how GitHub counts it.- Requests are retried three times on network and server errors.
Source and license
- Source: the GitHub REST API (
https://api.github.com, version2022-11-28) and, for manifests,https://raw.githubusercontent.com. Every row carriessourceandsourceLicense. - GitHub's terms: the API is used under Section H (API Terms) of the GitHub Terms of Service, which forbids excessive request rates and sharing tokens to exceed the rate limits. The actor stays within the limits, identifies itself with a User-Agent and stops when the limit is reached. GitHub's Acceptable Use Policies (Section 7, Information Usage Restrictions) forbid using information from GitHub for spamming or for selling personal information, for example to recruiters. This actor returns no personal information. You are responsible for using the data in line with GitHub's terms and the GitHub Privacy Statement.
- The repositories: metadata such as stars and release dates is facts about a repository. The code and files in each repository are under that repository's own license, shown in the
licensefield. The actor reads manifests only to count and name packages; it does not copy code.
This actor is not affiliated with or endorsed by GitHub.
Pricing
Pay per event: 0.004 USD per repository returned (event repository), which is 4.00 USD per 1 000 repositories. Platform usage is included, so you pay only per result. Error rows, unchanged repositories skipped by onlyNew and runs without results are not charged. maxResults caps how many repositories, and therefore how much, a run can charge.
Rows are delivered only after they have been charged. If you set a maximum cost per run (maxTotalChargeUsd), the run stops there and its status message says how many rows were delivered.
Contact
Built and maintained by Nightwave AB. Questions, bugs and feature requests: kontakt@nightwave.se
På svenska
Actorn hämtar fakta om GitHub-repon via GitHubs officiella API och ger en rad per repo: stjärnor, forks, ämnen, licens, språk med andel, senaste commit, senaste release och antal releaser, samt paketnamn, version och antal beroenden från package.json, pyproject.toml eller go.mod.
- Indata: en lista med repon (
owner/nameeller länk), en GitHub-sökning (till exempeltopic:web-scraping stars:>500) eller båda. Standard är 50 repon. Tom input läserapify/crawleeochapify/apify-sdk-jssom exempel. - Inga personuppgifter: actorn läser inga användarprofiler, e-postadresser eller namn på dem som gjort commits.
- Token är valfri. Utan token tillåter GitHub 60 anrop i timmen, vilket räcker till cirka 14 namngivna repon eller 19 från en sökning per körning. När gränsen nås avslutas körningen lugnt med en felrad per repo som inte lästes, och felrader debiteras inte. En fine-grained token med bara läsrätt till publika repon höjer gränsen till 5 000 anrop i timmen.
- Bevakning: med
onlyNewkommer actorn ihåg vad den har levererat (key-value storenightwave-state-github-repo-intelligence) och levererar bara repon med en ny push eller en ny release sedan förra körningen. Lägg den på ett dagligt schema i Apify under Schedules, se avsnittet "Monitoring and scheduling". - Källa: GitHub REST API enligt GitHubs användarvillkor (avsnitt H, API Terms). Koden i varje repo har sin egen licens, som står i fältet
license. - Pris: 0,004 USD per repo (4,00 USD per 1 000). Plattformsanvändningen ingår.
- Kontakt: kontakt@nightwave.se