GitHub Repository Scraper - Repos, Issues & Pull Requests avatar

GitHub Repository Scraper - Repos, Issues & Pull Requests

Pricing

Pay per usage

Go to Apify Store
GitHub Repository Scraper - Repos, Issues & Pull Requests

GitHub Repository Scraper - Repos, Issues & Pull Requests

Scrape GitHub repositories, issues, pull requests and contributors by search or repository link: stars, forks, topics, language, license, dates and authors. No GitHub token needed. Export to CSV, JSON or Excel, schedule runs or use the API.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Automly

Automly

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What is GitHub Repository Scraper?

GitHub Repository Scraper is a tool that lets you scrape repositories, open issues, open pull requests and contributor profiles from GitHub: stars, forks, open issue count, language, license, topics, issue and pull request titles, labels and authors, plus contributor names, companies, locations and public emails. Add a GitHub search or a list of repository links, click Start, and download the data as Excel, CSV or JSON.

  • ๐Ÿ”Ž Search or list: find repositories with GitHub search (language:python stars:>1000) or paste the exact repositories you want
  • ๐Ÿงฉ Four record types: repositories, issues, pull requests and contributors in one dataset, each row marked with its type
  • ๐Ÿ”“ No login: no GitHub account or key needed
  • ๐Ÿ’ธ Free to try: the $5 of free usage every Apify account gets each month covers about 10,000 repositories

What can GitHub Repository Scraper do?

  • Scrape GitHub repositories by search query, language, stars or topic
  • Scrape repository details from a list of GitHub URLs or owner/repo names
  • Export GitHub stars, forks, open issues, license and topics to Excel or CSV
  • Scrape open issues from GitHub repositories with labels, authors and comment counts
  • Scrape open pull requests from GitHub with merged and draft status
  • Get GitHub contributor profiles with company, location, bio and public email
  • Set how many issues, pull requests and contributors you want per repository
  • Export to Excel, CSV, JSON, HTML or XML, or send the data to Google Sheets, Make, Zapier and more

What data can you extract from GitHub repositories?

โญ Stars๐Ÿด Forks๐Ÿž Open issue count
๐Ÿ’ป Primary language๐Ÿ“œ License๐Ÿท๏ธ Topics
๐Ÿ“ Description๐Ÿ“… Created and updated dates๐Ÿ”— Repository URL
๐Ÿ› Issue title, state, labels๐Ÿ”€ Pull request merged and draft status๐Ÿ’ฌ Comment count
๐Ÿ‘ค Contributor name and username๐Ÿข Company and location๐Ÿ“ง Public email and blog

How to scrape GitHub repositories

  1. Create a free Apify account (no credit card needed).
  2. Open GitHub Repository Scraper.
  3. Type a Search Query such as language:typescript stars:>5000, or paste Repository URLs.
  4. Tick Extract Issues, Extract Pull Requests or Extract Contributors if you want them, set Maximum Results, and click Start.
  5. When the run finishes, download your data as Excel, CSV, JSON, HTML or XML.

How much does it cost to scrape GitHub repositories?

From 15 October 2026 you pay $0.50 per 1,000 repositories, $0.50 per 1,000 issues and $0.50 per 1,000 pull requests, on every Apify plan. Apify platform usage is included in these prices, so there is nothing else to pay for the run. Until then, see the Pricing tab for the current terms.

For example, from 15 October 2026 a run that returns 1,000 repositories and 1,000 of their issues costs $1.00 on the Free plan, and the same on every other plan.

Every Apify account gets $5 of free usage each month. See the Pricing tab for details.

โฌ‡๏ธ Input

Use either a search or a list of repositories. When Repository URLs has entries, the search is not used.

SettingWhat it does
Search QueryA GitHub search, e.g. language:python stars:>1000 or topic:web-scraping
Repository URLsRepository links or owner/repo names to scrape directly
Extract IssuesAlso get each repository's open issues (off by default)
Extract Pull RequestsAlso get each repository's open pull requests (off by default)
Extract ContributorsAlso get each repository's contributor profiles (off by default)
Maximum ResultsTotal records to return across all record types (default 100, up to 1,000)
Max Issues Per RepoOpen issues per repository (default 30)
Max Pull Requests Per RepoOpen pull requests per repository (default 30)
Max Contributors Per RepoContributors per repository (default 30)
Your own GitHub keyOptional and not needed. Leave it empty; the actor works without it

Example: popular TypeScript repositories with their issues and contributors.

{
"searchQuery": "language:typescript stars:>5000",
"extractIssues": true,
"extractUsers": true,
"maxResults": 50,
"maxIssuesPerRepo": 10,
"maxUsersPerRepo": 10
}

Example: one repository and its latest open issues.

{
"repoUrls": ["https://github.com/psf/requests"],
"extractIssues": true,
"maxIssuesPerRepo": 10,
"maxResults": 20
}

โฌ†๏ธ Output

You get one row per repository, issue, pull request or contributor. The type field tells them apart (repository, issue, pullRequest, user), and issue, pull request and contributor rows carry the repository they belong to. You can view the rows as a table in Apify Console or download them as Excel, CSV, JSON, HTML or XML.

A repository row:

{
"type": "repository",
"url": "https://github.com/psf/requests",
"owner": "psf",
"name": "requests",
"fullName": "psf/requests",
"description": "A simple, yet elegant, HTTP library.",
"stars": 54376,
"forks": 11508,
"openIssues": 242,
"language": "Python",
"license": "Apache-2.0",
"createdAt": "2011-02-13T18:38:17Z",
"updatedAt": "2026-10-02T05:30:07Z",
"topics": ["client", "cookies", "forhumans", "http", "humans", "python", "python-requests", "requests"]
}

An issue row from the same run:

{
"type": "issue",
"repository": "psf/requests",
"url": "https://github.com/psf/requests/issues/7631",
"number": 7631,
"title": "TestTimeout abandons three delay/10 requests and later tests wait out the server",
"state": "open",
"author": "quotentiroler",
"labels": [],
"createdAt": "2026-09-30T13:54:45Z",
"updatedAt": "2026-10-01T03:19:54Z",
"comments": 3
}

Repository fields

FieldDescription
typerepository
urlGitHub repository URL
ownerRepository owner
nameRepository name
fullNameowner/name
descriptionRepository description
starsStargazer count
forksFork count
openIssuesOpen issue count
languagePrimary language
licenseSPDX license identifier
createdAtCreation date and time (ISO 8601)
updatedAtLast update date and time (ISO 8601)
topicsRepository topics

Issue fields

FieldDescription
typeissue
repositoryParent repository
urlIssue URL
numberIssue number
titleIssue title
stateIssue state
authorAuthor username
labelsLabel names
createdAtCreation date and time (ISO 8601)
updatedAtLast update date and time (ISO 8601)
commentsComment count

Pull request fields

FieldDescription
typepullRequest
repositoryParent repository
urlPull request URL
numberPull request number
titlePull request title
statePull request state
authorAuthor username
createdAtCreation date and time (ISO 8601)
updatedAtLast update date and time (ISO 8601)
mergedMerged status
draftDraft status

Contributor fields

FieldDescription
typeuser
repositorySource repository
urlProfile URL
usernameGitHub username
nameDisplay name
companyCompany
blogBlog URL
locationLocation
emailPublic email
bioBio
publicReposPublic repository count
followersFollower count
followingFollowing count
createdAtAccount creation date and time (ISO 8601)

How can I use GitHub repository data?

  • Developer lead generation: find repositories by language, stars or topic and list the people who contribute to them, with company, location and public email
  • Competitive research: track stars, forks, open issues and pull requests across competitor and alternative projects
  • Talent sourcing: build lists of active contributors to the projects your team uses
  • Open-source health checks: compare activity, licenses and open issue counts before you adopt a library
  • Datasets for search and AI apps: feed repository descriptions, topics and issue titles into your own search or analysis tools

Scrape more GitHub and developer data

ActorWhat it gets
GitHub Issues Scraper - Issues & Pull Requests MonitorNew, updated, closed and merged issues and pull requests across repos, orgs and users, on a schedule
GitHub Users Scraper - Developers by Location & CompanyGitHub users and organizations by keyword, location, company and follower count
GitHub Code Search APIPublic code files that match your search, with repository details
Stack Overflow Scraper - Questions, Tags & ScoresStack Overflow and Stack Exchange questions with tags, scores and answer counts
npm Scraper - Package Metadata, Versions & Maintainersnpm package details, versions and maintainers
PyPI Scraper - Python Package Versions & DownloadsPython package metadata, versions and project links

โ“FAQ

How much does it cost to scrape GitHub repositories?

From 15 October 2026 it costs $0.50 per 1,000 repositories, issues or pull requests, with platform usage included. See the cost section above and the Pricing tab.

Up to 1,000. GitHub search shows at most 1,000 results per query, and Maximum Results caps the total number of records in a run (up to 1,000, counting repositories, issues, pull requests and contributors together). To cover more, split one broad search into narrower ones, for example by star range or language.

Do I need a GitHub account?

No. The actor works without a GitHub account, login or key.

Can I scrape private repositories?

No. It only collects public data that anyone can see on GitHub.

Why are some contributor emails empty?

GitHub only shows an email when the user has chosen to make it public. Many contributors keep theirs private, so the email field is empty for them.

Does it get closed issues and pull requests?

No. It gets open issues and open pull requests. To track closed and merged activity over time, use GitHub Issues Scraper - Issues & Pull Requests Monitor.

Is the data real-time?

Yes. The data reflects GitHub at the time of the run.

Can I connect GitHub Repository Scraper to other tools or AI agents?

Yes. Start runs and download results with the Apify API or the Python and JavaScript clients, send results to Make, Zapier, n8n, Google Sheets or Slack with Apify integrations and webhooks, or let AI assistants and agents run it through the Apify MCP server at mcp.apify.com.

This actor only collects data that is public on GitHub: repositories, issues, pull requests and the profile details users chose to show. Profiles can still contain personal data, such as names, locations and public emails, which may be protected by laws such as GDPR. Only scrape it for a legitimate reason and ask a lawyer if you are unsure. You can read more in Is web scraping legal?

GitHub Repository Scraper is an independent tool. It is not affiliated with, endorsed by or sponsored by GitHub.

Something isn't working?

Open the Issues tab and tell us your input and what you expected.

โญ Your feedback

Have an idea or found a problem? Tell us on the Issues tab. If GitHub Repository Scraper saved you time, a short review on the Reviews tab helps other people find it.