GitHub Scraper - Repos, Stars, Forks and Users avatar

GitHub Scraper - Repos, Stars, Forks and Users

Pricing

from $0.90 / 1,000 item returneds

Go to Apify Store
GitHub Scraper - Repos, Stars, Forks and Users

GitHub Scraper - Repos, Stars, Forks and Users

Search GitHub for repos and users. Use the same query you type in GitHub's search box, like `language:rust stars:>10000`. Repo rows give stars, forks, issues, language, topics, licence and dates. User rows add bio, company and followers. Any query stops at 1,000 results. $0.90 per 1,000 rows.

Pricing

from $0.90 / 1,000 item returneds

Rating

5.0

(2)

Developer

Dami's Studio

Dami's Studio

Maintained by Community

Actor stats

1

Bookmarked

42

Total users

15

Monthly active users

20 hours ago

Last modified

Share

GitHub Scraper: repositories and users from a GitHub search query

Type the query you would type into GitHub's own search box, language:rust stars:>10000, and you get one row per result. For repositories that is stars, forks, open issues, language, topics, licence and the three dates. For users it is the profile: name, bio, company, location, followers and public repo count.

The cap is GitHub's, not ours: any search stops at 1,000 results, however many matches it reports. Split a big job by star range or by date and run the pieces.

Worth knowing before you trust an empty result. GitHub answers a query with a broken qualifier, and a request with a bad token, in a way this actor currently reports as "no matches". So if a search you expect results from comes back with nothing, check the qualifier spelling and the token before concluding the answer is really empty.

InputOne GitHub search query, in GitHub's own syntax
OutputOne row per repository or user
Ceiling1,000 results per query, which is GitHub's own limit
Account neededNone. Your own GitHub token is optional and lifts the rate limits a long way
Price$0.90 per 1,000 rows, which is $0.0009 each, flat on every plan

🐙 What GitHub Scraper does

It runs your query against GitHub's search, pages through the results 100 at a time, and writes a row for each one. Repository searches sort by stars, forks, recent update or best match. User searches use GitHub's own relevance ranking and ignore the sort setting, because GitHub does.

Every qualifier GitHub documents works, since the query goes through as you wrote it: topic:cli created:>2023-01-01, location:berlin followers:>500, fullstack in:bio.

A user row is the search result plus a second call for the profile, so you get the bio, company and follower count rather than just a login.

Duplicates across pages are dropped before anything is written, so a repository appearing twice in GitHub's paging does not cost you twice.

Add your own GitHub token and the run gets GitHub's signed-in allowance, 5,000 requests an hour and 30 searches a minute, instead of the small signed-out one. No special scopes: public data only. Bigger jobs need it to finish.

📥 What you give it

{
"query": "language:python stars:>5000 web framework",
"type": "repositories",
"sort": "stars",
"maxItems": 200
}
FieldDefaultWhat it is
querynoneGitHub search syntax, exactly as the site takes it. The form arrives with an example in it.
typerepositoriesrepositories or users. Users and organisations both come back under users.
sortstarsstars, forks, updated or best-match. Repository searches only.
maxItems1001 to 1,000. GitHub stops at 1,000 for any single query.
githubTokennoneYour own personal access token, marked secret. No scopes needed for public data. Recommended for anything past a small run.
notionConnector, notionParentIdnoneOptional. Also write each row into your connected Notion workspace.
proxyConfigurationnoneOptional and usually unnecessary.

📤 What you get back

A real repository row from a real run:

{
"ok": true,
"fullName": "fastapi/fastapi",
"name": "fastapi",
"owner": "fastapi",
"url": "https://github.com/fastapi/fastapi",
"description": "FastAPI framework, high performance, easy to learn, fast to code, ready for production",
"stars": 102502,
"forks": 9921,
"openIssues": 82,
"language": "Python",
"topics": ["api", "async", "asyncio", "fastapi", "framework", "json", "openapi", "python", "rest", "swagger"],
"license": "MIT",
"homepage": "https://fastapi.tiangolo.com/",
"defaultBranch": "master",
"createdAt": "2018-12-08T08:21:47Z",
"updatedAt": "2026-09-21T12:45:53Z",
"pushedAt": "2026-09-18T21:24:37Z"
}

Repository rows:

FieldWhat it is
fullName, name, owner, urlThe repo and who owns it. fullName is your key.
stars, forks, openIssuesNumbers at read time, not running totals.
language, topics, licenseThe main language GitHub detects, the topic tags, and the licence as its SPDX id.
homepage, defaultBranchAs set on the repo. Null when there is none.
createdAt, updatedAt, pushedAtpushedAt is the one that tells you whether a project is alive. updatedAt moves on a star, too.

User rows, from a type: users run:

FieldWhat it is
login, url, type, idThe account. type separates a person from an organisation.
name, bio, company, location, blogProfile text, as filled in. Plenty of accounts leave these empty.
followers, publicRepos, createdAtFollower count, public repo count, and when the account was made.

🧾 Reading the output

Data rows carry ok: true. Notes carry ok: false and an errorCode, and are never charged.

RowHow to spot itCharged
A repository or userok is trueyes
A diagnosticok is false, and errorCode says what happenedno
CodeWhat it means
BAD_INPUTThe query was empty.
NO_RESULTSGitHub returned no matches. Check the qualifier spelling and your token before believing it, since a rejected query currently arrives here too.
RATE_LIMITEDGitHub's limit was hit. rateLimitResetsAt says when it clears. Add a token, or ask for fewer items.
BLOCKEDGitHub refused the request. Usually its short-term limit on bursts, so wait a minute and re-run.
NOT_FOUNDThe thing being asked for is not there.
SERVER_ERRORGitHub's own error. Worth one re-run.
NETWORKGitHub was unreachable or answered badly.

The default table view is built for repositories, so a users run looks blank in it. Open a row, or export the dataset, and the profile fields are all there.

▶️ How to run it

  1. Open GitHub Scraper and click Try for free.
  2. Put your query into Search query. Build it on GitHub first if you are not sure of it, then paste it across.
  3. Pick Search type: repositories or users.
  4. Set Max items, up to 1,000. For anything past a hundred or so, paste a GitHub token.
  5. Click Start, then download the dataset as JSON, CSV or Excel, or read it from the API.

💰 How much does it cost?

$0.90 per 1,000 rows, which is $0.0009 each. Flat on every Apify plan, no volume tiers.

You pay per result row, repository or user. Diagnostics are not charged, and neither are duplicates that GitHub's paging returned twice, since they are dropped before anything is written.

A thousand repositories, which is the most any single query can give you, is ninety cents.

💡 What people use it for

  • Finding maintainers to talk to: location:berlin language:go followers:>200 as a user search.
  • Tracking a topic over time. The same query weekly, keyed on fullName, shows what is growing.
  • Competitive work on open source: who forked what, which projects stopped being pushed to.
  • Licence audits across a topic, using the SPDX id on every row.
  • Feeding a Notion database of interesting repos straight from the run.

🚧 What it does not do

  • 1,000 results per query. GitHub's cap. Split by stars:1000..5000, by created: window, or by language.
  • No repo contents. No README, no file tree, no issues, no pull requests, no contributor list.
  • No private data. Public repositories and public profiles only, whatever token you paste.
  • sort does nothing on a user search. GitHub ranks those itself.
  • A rejected query reads as an empty one. A misspelt qualifier or a bad token comes back as "no matches" rather than as an error.
  • A rate limit part-way through loses that run's rows. You get the note rather than a partial page, so use a token for long jobs.
  • The odd user row comes back thin, with only the login and URL, when GitHub's profile call did not answer for that account.
  • Counts are a snapshot. Stars move.

🧭 Which developer data scraper do you need?

If you wantUse
GitHub repositories or users from a searchThis one
Questions, tags and scores from Stack OverflowStack Overflow Scraper
npm and PyPI package metadata and downloadsnpm + PyPI Package Scraper
Structured facts and entity claimsWikidata Scraper

❓ Questions people ask

Do I need a GitHub token? Not for a small run. For anything bigger, yes, because the signed-out allowance is tight and a run that hits it stops.

Is my token safe? It goes in a secret input field, it is only used against GitHub's own API, and it never appears in the dataset or the log. A read-only token with no scopes is enough.

Why did my query return nothing? Check the qualifier first. stars:>1000 works, star:>1000 does not, and a rejected query looks like an empty one here.

How do I get more than 1,000 repos? Run several queries with narrower ranges. Star bands and creation-date windows split most jobs cleanly.

Can I search organisations? Yes, they come back in a users search with type telling you which is which.

Is scraping GitHub legal? This uses GitHub's own public search API and returns public data. Profiles are personal data under GDPR and similar laws, so have a reason for holding them. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the exact query and the run ID. If a diagnostic row landed, its errorCode and hint usually name the reason already. Never paste your token into an issue.