GitHub Repository Scraper - Repos, Issues & Pull Requests
Pricing
Pay per usage
GitHub Repository Scraper - Repos, Issues & Pull Requests
Scrape GitHub repositories, issues, pull requests and contributors by search or repository link: stars, forks, topics, language, license, dates and authors. No GitHub token needed. Export to CSV, JSON or Excel, schedule runs or use the API.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Automly
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
What is GitHub Repository Scraper?
GitHub Repository Scraper is a tool that lets you scrape repositories, open issues, open pull requests and contributor profiles from GitHub: stars, forks, open issue count, language, license, topics, issue and pull request titles, labels and authors, plus contributor names, companies, locations and public emails. Add a GitHub search or a list of repository links, click Start, and download the data as Excel, CSV or JSON.
- ๐ Search or list: find repositories with GitHub search (
language:python stars:>1000) or paste the exact repositories you want - ๐งฉ Four record types: repositories, issues, pull requests and contributors in one dataset, each row marked with its
type - ๐ No login: no GitHub account or key needed
- ๐ธ Free to try: the $5 of free usage every Apify account gets each month covers about 10,000 repositories
What can GitHub Repository Scraper do?
- Scrape GitHub repositories by search query, language, stars or topic
- Scrape repository details from a list of GitHub URLs or
owner/reponames - Export GitHub stars, forks, open issues, license and topics to Excel or CSV
- Scrape open issues from GitHub repositories with labels, authors and comment counts
- Scrape open pull requests from GitHub with merged and draft status
- Get GitHub contributor profiles with company, location, bio and public email
- Set how many issues, pull requests and contributors you want per repository
- Export to Excel, CSV, JSON, HTML or XML, or send the data to Google Sheets, Make, Zapier and more
What data can you extract from GitHub repositories?
| โญ Stars | ๐ด Forks | ๐ Open issue count |
| ๐ป Primary language | ๐ License | ๐ท๏ธ Topics |
| ๐ Description | ๐ Created and updated dates | ๐ Repository URL |
| ๐ Issue title, state, labels | ๐ Pull request merged and draft status | ๐ฌ Comment count |
| ๐ค Contributor name and username | ๐ข Company and location | ๐ง Public email and blog |
How to scrape GitHub repositories
- Create a free Apify account (no credit card needed).
- Open GitHub Repository Scraper.
- Type a Search Query such as
language:typescript stars:>5000, or paste Repository URLs. - Tick Extract Issues, Extract Pull Requests or Extract Contributors if you want them, set Maximum Results, and click Start.
- When the run finishes, download your data as Excel, CSV, JSON, HTML or XML.
How much does it cost to scrape GitHub repositories?
From 15 October 2026 you pay $0.50 per 1,000 repositories, $0.50 per 1,000 issues and $0.50 per 1,000 pull requests, on every Apify plan. Apify platform usage is included in these prices, so there is nothing else to pay for the run. Until then, see the Pricing tab for the current terms.
For example, from 15 October 2026 a run that returns 1,000 repositories and 1,000 of their issues costs $1.00 on the Free plan, and the same on every other plan.
Every Apify account gets $5 of free usage each month. See the Pricing tab for details.
โฌ๏ธ Input
Use either a search or a list of repositories. When Repository URLs has entries, the search is not used.
| Setting | What it does |
|---|---|
| Search Query | A GitHub search, e.g. language:python stars:>1000 or topic:web-scraping |
| Repository URLs | Repository links or owner/repo names to scrape directly |
| Extract Issues | Also get each repository's open issues (off by default) |
| Extract Pull Requests | Also get each repository's open pull requests (off by default) |
| Extract Contributors | Also get each repository's contributor profiles (off by default) |
| Maximum Results | Total records to return across all record types (default 100, up to 1,000) |
| Max Issues Per Repo | Open issues per repository (default 30) |
| Max Pull Requests Per Repo | Open pull requests per repository (default 30) |
| Max Contributors Per Repo | Contributors per repository (default 30) |
| Your own GitHub key | Optional and not needed. Leave it empty; the actor works without it |
Example: popular TypeScript repositories with their issues and contributors.
{"searchQuery": "language:typescript stars:>5000","extractIssues": true,"extractUsers": true,"maxResults": 50,"maxIssuesPerRepo": 10,"maxUsersPerRepo": 10}
Example: one repository and its latest open issues.
{"repoUrls": ["https://github.com/psf/requests"],"extractIssues": true,"maxIssuesPerRepo": 10,"maxResults": 20}
โฌ๏ธ Output
You get one row per repository, issue, pull request or contributor. The type field tells them apart (repository, issue, pullRequest, user), and issue, pull request and contributor rows carry the repository they belong to. You can view the rows as a table in Apify Console or download them as Excel, CSV, JSON, HTML or XML.
A repository row:
{"type": "repository","url": "https://github.com/psf/requests","owner": "psf","name": "requests","fullName": "psf/requests","description": "A simple, yet elegant, HTTP library.","stars": 54376,"forks": 11508,"openIssues": 242,"language": "Python","license": "Apache-2.0","createdAt": "2011-02-13T18:38:17Z","updatedAt": "2026-10-02T05:30:07Z","topics": ["client", "cookies", "forhumans", "http", "humans", "python", "python-requests", "requests"]}
An issue row from the same run:
{"type": "issue","repository": "psf/requests","url": "https://github.com/psf/requests/issues/7631","number": 7631,"title": "TestTimeout abandons three delay/10 requests and later tests wait out the server","state": "open","author": "quotentiroler","labels": [],"createdAt": "2026-09-30T13:54:45Z","updatedAt": "2026-10-01T03:19:54Z","comments": 3}
Repository fields
| Field | Description |
|---|---|
| type | repository |
| url | GitHub repository URL |
| owner | Repository owner |
| name | Repository name |
| fullName | owner/name |
| description | Repository description |
| stars | Stargazer count |
| forks | Fork count |
| openIssues | Open issue count |
| language | Primary language |
| license | SPDX license identifier |
| createdAt | Creation date and time (ISO 8601) |
| updatedAt | Last update date and time (ISO 8601) |
| topics | Repository topics |
Issue fields
| Field | Description |
|---|---|
| type | issue |
| repository | Parent repository |
| url | Issue URL |
| number | Issue number |
| title | Issue title |
| state | Issue state |
| author | Author username |
| labels | Label names |
| createdAt | Creation date and time (ISO 8601) |
| updatedAt | Last update date and time (ISO 8601) |
| comments | Comment count |
Pull request fields
| Field | Description |
|---|---|
| type | pullRequest |
| repository | Parent repository |
| url | Pull request URL |
| number | Pull request number |
| title | Pull request title |
| state | Pull request state |
| author | Author username |
| createdAt | Creation date and time (ISO 8601) |
| updatedAt | Last update date and time (ISO 8601) |
| merged | Merged status |
| draft | Draft status |
Contributor fields
| Field | Description |
|---|---|
| type | user |
| repository | Source repository |
| url | Profile URL |
| username | GitHub username |
| name | Display name |
| company | Company |
| blog | Blog URL |
| location | Location |
| Public email | |
| bio | Bio |
| publicRepos | Public repository count |
| followers | Follower count |
| following | Following count |
| createdAt | Account creation date and time (ISO 8601) |
How can I use GitHub repository data?
- Developer lead generation: find repositories by language, stars or topic and list the people who contribute to them, with company, location and public email
- Competitive research: track stars, forks, open issues and pull requests across competitor and alternative projects
- Talent sourcing: build lists of active contributors to the projects your team uses
- Open-source health checks: compare activity, licenses and open issue counts before you adopt a library
- Datasets for search and AI apps: feed repository descriptions, topics and issue titles into your own search or analysis tools
Scrape more GitHub and developer data
| Actor | What it gets |
|---|---|
| GitHub Issues Scraper - Issues & Pull Requests Monitor | New, updated, closed and merged issues and pull requests across repos, orgs and users, on a schedule |
| GitHub Users Scraper - Developers by Location & Company | GitHub users and organizations by keyword, location, company and follower count |
| GitHub Code Search API | Public code files that match your search, with repository details |
| Stack Overflow Scraper - Questions, Tags & Scores | Stack Overflow and Stack Exchange questions with tags, scores and answer counts |
| npm Scraper - Package Metadata, Versions & Maintainers | npm package details, versions and maintainers |
| PyPI Scraper - Python Package Versions & Downloads | Python package metadata, versions and project links |
โFAQ
How much does it cost to scrape GitHub repositories?
From 15 October 2026 it costs $0.50 per 1,000 repositories, issues or pull requests, with platform usage included. See the cost section above and the Pricing tab.
How many repositories can I get from one search?
Up to 1,000. GitHub search shows at most 1,000 results per query, and Maximum Results caps the total number of records in a run (up to 1,000, counting repositories, issues, pull requests and contributors together). To cover more, split one broad search into narrower ones, for example by star range or language.
Do I need a GitHub account?
No. The actor works without a GitHub account, login or key.
Can I scrape private repositories?
No. It only collects public data that anyone can see on GitHub.
Why are some contributor emails empty?
GitHub only shows an email when the user has chosen to make it public. Many contributors keep theirs private, so the email field is empty for them.
Does it get closed issues and pull requests?
No. It gets open issues and open pull requests. To track closed and merged activity over time, use GitHub Issues Scraper - Issues & Pull Requests Monitor.
Is the data real-time?
Yes. The data reflects GitHub at the time of the run.
Can I connect GitHub Repository Scraper to other tools or AI agents?
Yes. Start runs and download results with the Apify API or the Python and JavaScript clients, send results to Make, Zapier, n8n, Google Sheets or Slack with Apify integrations and webhooks, or let AI assistants and agents run it through the Apify MCP server at mcp.apify.com.
Is it legal to scrape GitHub?
This actor only collects data that is public on GitHub: repositories, issues, pull requests and the profile details users chose to show. Profiles can still contain personal data, such as names, locations and public emails, which may be protected by laws such as GDPR. Only scrape it for a legitimate reason and ask a lawyer if you are unsure. You can read more in Is web scraping legal?
GitHub Repository Scraper is an independent tool. It is not affiliated with, endorsed by or sponsored by GitHub.
Something isn't working?
Open the Issues tab and tell us your input and what you expected.
โญ Your feedback
Have an idea or found a problem? Tell us on the Issues tab. If GitHub Repository Scraper saved you time, a short review on the Reviews tab helps other people find it.