GitHub Scraper — Repos, Issues, PRs & Code avatar

GitHub Scraper — Repos, Issues, PRs & Code

Pricing

Pay per event + usage

Go to Apify Store
GitHub Scraper — Repos, Issues, PRs & Code

GitHub Scraper — Repos, Issues, PRs & Code

Scrape GitHub deeply — repos, issues, PRs, code search, contributors, releases, READMEs, commits, users, trending. 11 modes in one actor for AI coding agents (Claude Code, Cursor, Copilot). Optional PAT for 5K req/hr. MCP-ready, flat JSON output.

Pricing

Pay per event + usage

Rating

0.0

(0)

Developer

Khadin Akbar

Khadin Akbar

Maintained by Community

Actor stats

0

Bookmarked

23

Total users

9

Monthly active users

14 days ago

Last modified

Share

GitHub Deep Scraper — Repos, Issues, PRs, Code Search, Commits

GitHub Deep Scraper is an Apify Actor for public GitHub data. It accepts one GitHub surface per run, such as a repository, search query, user login, or trending filter, and returns one flat record per result. Each record includes mode, type, url, and scrapedAt, plus mode-specific GitHub fields that support repo analysis, issue review, PR inspection, code search, contributor lookup, release tracking, README retrieval, commit history review, user profiling, and trending exploration.

Best fit and connected workflows

This Actor fits workflows that need one GitHub scraping tool with multiple routes:

  • Repository research: use repo when a workflow starts from one public repository and needs a full metadata snapshot.
  • Issue and PR review: use issues or prs when a workflow starts from a repo and needs lists, labels, comments, reviews, or review threads.
  • Code discovery: use code-search when a workflow starts from a text query and needs code matches across GitHub.
  • Contributor and release analysis: use contributors or releases when the workflow needs contributor profiles or release history from a repo.
  • Documentation extraction: use readme when the workflow needs the repository README as markdown and rendered text.
  • Commit inspection: use commits when the workflow needs commit history, and optionally file diffs.
  • Profile review: use user when the workflow starts from a GitHub login and needs the profile plus repositories.
  • Trend scanning: use trending when the workflow needs trending repositories by language and timeframe.

Practical scenario

A developer advocate named Mina starts with apify/actors-mcp-server and wants to prepare a short update for her team. She runs the Actor in prs mode with state: "closed" and includeReviews: true. The returned dataset record includes the PR url, type: "pr", and scrapedAt, plus PR-specific fields for review state and review comments. Mina uses the review state and PR URL to pick one merged change for a deeper read, then shares the PR link in her team note.

Input fields

FieldTypeUsed inDescription
modestringall modesSelects one GitHub surface to scrape. One mode per run.
repostringrepo, issues, prs, contributors, releases, readme, commitsGitHub repository in owner/name format, or a full GitHub repo URL.
querystringrepo-search, code-searchGitHub search query with qualifiers such as language:, stars:, org:, path:, and extension:.
userstringuserGitHub user or organization login.
languagestringtrendingLanguage filter for trending mode.
timeframestringtrendingTrending window: daily, weekly, or monthly.
statestringissues, prsIssue or PR state filter: open, closed, or all.
sincestringissues, prs, commitsISO 8601 date for returning items updated or created at or after the given point.
maxResultsintegerall modesMaximum number of records to return.
includeCommentsbooleanissues, prsIncludes comments for each issue or PR.
includeReviewsbooleanprsIncludes reviews and review comments for each PR.
includeFilesbooleancommitsIncludes file diffs for each commit.

Focused JSON input example

{
"mode": "prs",
"repo": "vercel/next.js",
"state": "closed",
"includeReviews": true,
"since": "2026-01-01",
"maxResults": 25
}

Output fields

The dataset returns one record per result.

FieldTypeDescription
modestringWhich scraping mode produced the record.
typestringRecord type such as repo, issue, pr, code-match, contributor, release, readme, commit, user, or trending-repo.
urlstring or nullCanonical GitHub URL for the entity.
scrapedAtstringISO 8601 UTC timestamp for when the record was fetched.

Illustrative JSON record

{
"mode": "repo",
"type": "repo",
"url": "https://github.com/facebook/react",
"scrapedAt": "2026-05-28T18:14:32Z"
}

How it works

This Actor uses GitHub's public REST API and GraphQL API surfaces for the selected mode, plus the public GitHub trending page for trending mode. The input schema defines 11 modes: repo, repo-search, issues, prs, code-search, contributors, releases, readme, commits, user, and trending. Repository-based modes normalize owner/name input, and full repository URLs are accepted for the repo field. The output dataset stays flat so downstream tools and agents can read it easily. The output contract also includes a GitHub rate-limit record in key-value storage, plus summary records for the run outcome.

Pricing

This Actor uses Apify Pay per event plus standard Apify platform usage. Open the live Pricing tab in Apify Console for the current event definitions and platform usage details before planning larger runs.

Charged events include:

  • apify-actor-start for the run start
  • result for simple returned records
  • deep-result for heavier returned records

For example, a run that returns twenty simple results is billed for one actor start event plus twenty result events, alongside Apify platform usage for the run.

Use with AI agents (MCP)

This Actor is Apify MCP-ready and usable through Apify MCP as khadinakbar/github-deep-scraper. It is a tool for retrieving structured GitHub data by mode, so an agent can request exactly the surface it needs and then read the dataset rows returned by the run.

Tool description: fetch public GitHub records for a chosen mode, then read the resulting dataset items and supporting run records such as rate-limit state.

Fetch the closed pull requests for vercel/next.js since the start of the year, include reviews, and return the dataset items with their GitHub URLs and scrape timestamps.

When interpreting results, treat each dataset row as one GitHub entity or search result. The shared fields identify the source mode, record type, canonical URL, and fetch time. For broader runs, use maxResults to bound how many rows the agent receives. For heavier modes and options, such as code search, PR reviews, PR comments, and commit file diffs, the deep-result event rate applies. The rate-limit record in key-value storage can help agents confirm whether a GitHub token is being applied and how many requests remain.

Apify API example

import { ApifyClient } from "apify-client";
const client = new ApifyClient({
token: process.env.APIFY_TOKEN,
});
const run = await client.actor("khadinakbar/github-deep-scraper").call({
mode: "repo",
repo: "facebook/react",
maxResults: 1,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

Best results and outcome guidance

Use repo when the workflow starts from one repository and needs a broad snapshot. Use repo-search when the workflow starts from a topic, stack, or qualifier query. Use issues and prs when the workflow needs discussion context, labels, reviews, or thread comments. Use commits when the workflow needs change history, and enable file diffs when reasoning about touched files matters. Use trending when the workflow needs a current list of popular repositories by language and timeframe. Set maxResults to the smallest number that covers the task so the run stays focused.

Continue the workflow

Design note

I found that the dataset contract is intentionally flat: every record includes mode, type, url, and scrapedAt, while the rest of the fields vary by mode.

FAQ

Which GitHub surface should I use for a full repository snapshot?
Use repo for one repository's full metadata.

Which mode is suited to finding repositories by topic or qualifier?
Use repo-search and pass a GitHub search query in the query field.

Which mode should I use for issue threads or PR conversations?
Use issues or prs, and turn on includeComments when you want thread context.

How do I inspect code across GitHub?
Use code-search with a search query. The live contract marks this mode as requiring the GITHUB_TOKEN environment variable.

How do I get a repository README?
Use readme with the repository name in owner/name format.

How do I get commit history with file diffs?
Use commits and set includeFiles: true when file-level reasoning matters.

How do I get trending repositories for a language?
Use trending with language and timeframe.

Responsible use

Use this Actor for public GitHub data and keep your usage aligned with GitHub's terms and applicable data-protection rules. Review the live Pricing tab before larger runs, especially when using result-heavy modes or options that add comments, reviews, or file diffs.