GitHub Scraper - Repos, Stars, Forks and Users
Pricing
from $0.90 / 1,000 item returneds
GitHub Scraper - Repos, Stars, Forks and Users
Search GitHub for repos and users. Use the same query you type in GitHub's search box, like `language:rust stars:>10000`. Repo rows give stars, forks, issues, language, topics, licence and dates. User rows add bio, company and followers. Any query stops at 1,000 results. $0.90 per 1,000 rows.
Pricing
from $0.90 / 1,000 item returneds
Rating
5.0
(2)
Developer
Dami's Studio
Maintained by CommunityActor stats
1
Bookmarked
42
Total users
15
Monthly active users
20 hours ago
Last modified
Share
GitHub Scraper: repositories and users from a GitHub search query
Type the query you would type into GitHub's own search box, language:rust stars:>10000, and you
get one row per result. For repositories that is stars, forks, open issues, language, topics,
licence and the three dates. For users it is the profile: name, bio, company, location, followers
and public repo count.
The cap is GitHub's, not ours: any search stops at 1,000 results, however many matches it reports. Split a big job by star range or by date and run the pieces.
Worth knowing before you trust an empty result. GitHub answers a query with a broken qualifier, and a request with a bad token, in a way this actor currently reports as "no matches". So if a search you expect results from comes back with nothing, check the qualifier spelling and the token before concluding the answer is really empty.
| Input | One GitHub search query, in GitHub's own syntax |
| Output | One row per repository or user |
| Ceiling | 1,000 results per query, which is GitHub's own limit |
| Account needed | None. Your own GitHub token is optional and lifts the rate limits a long way |
| Price | $0.90 per 1,000 rows, which is $0.0009 each, flat on every plan |
🐙 What GitHub Scraper does
It runs your query against GitHub's search, pages through the results 100 at a time, and writes a row for each one. Repository searches sort by stars, forks, recent update or best match. User searches use GitHub's own relevance ranking and ignore the sort setting, because GitHub does.
Every qualifier GitHub documents works, since the query goes through as you wrote it:
topic:cli created:>2023-01-01, location:berlin followers:>500, fullstack in:bio.
A user row is the search result plus a second call for the profile, so you get the bio, company and follower count rather than just a login.
Duplicates across pages are dropped before anything is written, so a repository appearing twice in GitHub's paging does not cost you twice.
Add your own GitHub token and the run gets GitHub's signed-in allowance, 5,000 requests an hour and 30 searches a minute, instead of the small signed-out one. No special scopes: public data only. Bigger jobs need it to finish.
📥 What you give it
{"query": "language:python stars:>5000 web framework","type": "repositories","sort": "stars","maxItems": 200}
| Field | Default | What it is |
|---|---|---|
query | none | GitHub search syntax, exactly as the site takes it. The form arrives with an example in it. |
type | repositories | repositories or users. Users and organisations both come back under users. |
sort | stars | stars, forks, updated or best-match. Repository searches only. |
maxItems | 100 | 1 to 1,000. GitHub stops at 1,000 for any single query. |
githubToken | none | Your own personal access token, marked secret. No scopes needed for public data. Recommended for anything past a small run. |
notionConnector, notionParentId | none | Optional. Also write each row into your connected Notion workspace. |
proxyConfiguration | none | Optional and usually unnecessary. |
📤 What you get back
A real repository row from a real run:
{"ok": true,"fullName": "fastapi/fastapi","name": "fastapi","owner": "fastapi","url": "https://github.com/fastapi/fastapi","description": "FastAPI framework, high performance, easy to learn, fast to code, ready for production","stars": 102502,"forks": 9921,"openIssues": 82,"language": "Python","topics": ["api", "async", "asyncio", "fastapi", "framework", "json", "openapi", "python", "rest", "swagger"],"license": "MIT","homepage": "https://fastapi.tiangolo.com/","defaultBranch": "master","createdAt": "2018-12-08T08:21:47Z","updatedAt": "2026-09-21T12:45:53Z","pushedAt": "2026-09-18T21:24:37Z"}
Repository rows:
| Field | What it is |
|---|---|
fullName, name, owner, url | The repo and who owns it. fullName is your key. |
stars, forks, openIssues | Numbers at read time, not running totals. |
language, topics, license | The main language GitHub detects, the topic tags, and the licence as its SPDX id. |
homepage, defaultBranch | As set on the repo. Null when there is none. |
createdAt, updatedAt, pushedAt | pushedAt is the one that tells you whether a project is alive. updatedAt moves on a star, too. |
User rows, from a type: users run:
| Field | What it is |
|---|---|
login, url, type, id | The account. type separates a person from an organisation. |
name, bio, company, location, blog | Profile text, as filled in. Plenty of accounts leave these empty. |
followers, publicRepos, createdAt | Follower count, public repo count, and when the account was made. |
🧾 Reading the output
Data rows carry ok: true. Notes carry ok: false and an errorCode, and are never charged.
| Row | How to spot it | Charged |
|---|---|---|
| A repository or user | ok is true | yes |
| A diagnostic | ok is false, and errorCode says what happened | no |
| Code | What it means |
|---|---|
BAD_INPUT | The query was empty. |
NO_RESULTS | GitHub returned no matches. Check the qualifier spelling and your token before believing it, since a rejected query currently arrives here too. |
RATE_LIMITED | GitHub's limit was hit. rateLimitResetsAt says when it clears. Add a token, or ask for fewer items. |
BLOCKED | GitHub refused the request. Usually its short-term limit on bursts, so wait a minute and re-run. |
NOT_FOUND | The thing being asked for is not there. |
SERVER_ERROR | GitHub's own error. Worth one re-run. |
NETWORK | GitHub was unreachable or answered badly. |
The default table view is built for repositories, so a users run looks blank in it. Open a row, or
export the dataset, and the profile fields are all there.
▶️ How to run it
- Open GitHub Scraper and click Try for free.
- Put your query into Search query. Build it on GitHub first if you are not sure of it, then paste it across.
- Pick Search type: repositories or users.
- Set Max items, up to 1,000. For anything past a hundred or so, paste a GitHub token.
- Click Start, then download the dataset as JSON, CSV or Excel, or read it from the API.
💰 How much does it cost?
$0.90 per 1,000 rows, which is $0.0009 each. Flat on every Apify plan, no volume tiers.
You pay per result row, repository or user. Diagnostics are not charged, and neither are duplicates that GitHub's paging returned twice, since they are dropped before anything is written.
A thousand repositories, which is the most any single query can give you, is ninety cents.
💡 What people use it for
- Finding maintainers to talk to:
location:berlin language:go followers:>200as a user search. - Tracking a topic over time. The same query weekly, keyed on
fullName, shows what is growing. - Competitive work on open source: who forked what, which projects stopped being pushed to.
- Licence audits across a topic, using the SPDX id on every row.
- Feeding a Notion database of interesting repos straight from the run.
🚧 What it does not do
- 1,000 results per query. GitHub's cap. Split by
stars:1000..5000, bycreated:window, or by language. - No repo contents. No README, no file tree, no issues, no pull requests, no contributor list.
- No private data. Public repositories and public profiles only, whatever token you paste.
sortdoes nothing on a user search. GitHub ranks those itself.- A rejected query reads as an empty one. A misspelt qualifier or a bad token comes back as "no matches" rather than as an error.
- A rate limit part-way through loses that run's rows. You get the note rather than a partial page, so use a token for long jobs.
- The odd user row comes back thin, with only the login and URL, when GitHub's profile call did not answer for that account.
- Counts are a snapshot. Stars move.
🧭 Which developer data scraper do you need?
| If you want | Use |
|---|---|
| GitHub repositories or users from a search | This one |
| Questions, tags and scores from Stack Overflow | Stack Overflow Scraper |
| npm and PyPI package metadata and downloads | npm + PyPI Package Scraper |
| Structured facts and entity claims | Wikidata Scraper |
❓ Questions people ask
Do I need a GitHub token? Not for a small run. For anything bigger, yes, because the signed-out allowance is tight and a run that hits it stops.
Is my token safe? It goes in a secret input field, it is only used against GitHub's own API, and it never appears in the dataset or the log. A read-only token with no scopes is enough.
Why did my query return nothing? Check the qualifier first. stars:>1000 works, star:>1000
does not, and a rejected query looks like an empty one here.
How do I get more than 1,000 repos? Run several queries with narrower ranges. Star bands and creation-date windows split most jobs cleanly.
Can I search organisations? Yes, they come back in a users search with type telling you
which is which.
Is scraping GitHub legal? This uses GitHub's own public search API and returns public data. Profiles are personal data under GDPR and similar laws, so have a reason for holding them. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the exact query and the run ID. If a diagnostic row
landed, its errorCode and hint usually name the reason already. Never paste your token into an
issue.