GitHub Contributor Email Scraper avatar

GitHub Contributor Email Scraper

Pricing

from $0.02 / 1,000 contributor email rows

Go to Apify Store
GitHub Contributor Email Scraper

GitHub Contributor Email Scraper

Export public GitHub commit-author emails, contributor identity fields, timestamps, and commit evidence from repositories or organizations.

Pricing

from $0.02 / 1,000 contributor email rows

Rating

0.0

(0)

Developer

Hanna Nosova

Hanna Nosova

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

9 days ago

Last modified

Share

Export public GitHub commit-author emails, contributor identity fields, timestamps, and commit evidence from repositories or organizations.

GitHub Contributor Email Scraper uses GitHub's official APIs and saves clean dataset rows for CSV, Excel, JSON, API, automation, and authorized research workflows. It supports tokenless REST for small public jobs and authenticated GraphQL for efficient, quota-aware commit-history queries.

What data does it export?

  • Public commit-author email and name
  • GitHub login and profile URL when GitHub links the commit to an account
  • Repository owner and name
  • Commit SHA and evidence URL
  • Commit timestamp and a short message snippet
  • Noreply-email status
  • Explicit-repository or organization-discovery provenance
  • REST or GraphQL route provenance

The Actor does not infer private emails, access private repositories, use login cookies, or bypass GitHub privacy settings.

Ready-to-run examples

View all ready-to-run examples.

Example input

{
"repositories": [
"apify/crawlee",
"apify/apify-sdk-js"
],
"maxCommitsPerRepo": 100,
"since": "2026-01-01",
"includeNoreplyEmails": false,
"dedupeByEmail": true,
"apiMode": "auto"
}

Use repositoryUrls when you have full GitHub URLs, repositories for owner/repo values, or organization to discover public repositories belonging to an organization.

Example output

{
"repository": "apify/crawlee",
"owner": "apify",
"repo": "crawlee",
"commitSha": "abc123",
"commitUrl": "https://github.com/apify/crawlee/commit/abc123",
"authorName": "Example Developer",
"authorEmail": "dev@example.com",
"isNoreplyEmail": false,
"githubLogin": "exampledev",
"profileUrl": "https://github.com/exampledev",
"committedAt": "2026-01-01T12:00:00Z",
"messageSnippet": "Improve retry handling",
"sourceType": "repository",
"scrapedAt": "2026-07-20T12:00:00Z",
"apiMode": "graphql"
}

Every dataset row keeps a commit URL and SHA so you can review the public source evidence.

Input settings

SettingAPI keyDescription
Repository URLsrepositoryUrlsUp to 100 public GitHub repository URLs.
Repository namesrepositoriesUp to 100 owner/repo values or full repository URLs.
Organization loginorganizationOptional GitHub organization login or profile URL.
Maximum organization repositoriesmaxRepositoriesPublic organization repositories to discover, from 1 to 100. Default: 20.
Maximum commits per repositorymaxCommitsPerRepoRecent commits inspected per repository, from 1 to 1,000. Default: 20.
Since datesinceOptional ISO date or timestamp lower bound.
Until dateuntilOptional ISO date or timestamp upper bound.
BranchbranchOptional branch name. Blank uses each repository's default branch.
Include noreply emailsincludeNoreplyEmailsKeep GitHub noreply addresses. Default: false.
Deduplicate by emaildedupeByEmailSave the first row for each email across the run. Default: true.
GitHub API modeapiModeauto, rest, or graphql. Default: auto.
Maximum active run secondsmaxRunSecondsOptional active-work cutoff from 1 to 270 seconds. Pending work is checkpointed.
GitHub tokengithubTokenOptional secret token for GraphQL and higher GitHub limits. Public-repository read access is sufficient.

At least one repository input or an organization is required. Invalid GitHub hosts, malformed repository names, invalid dates, and unsafe branch names are rejected before paid or network work.

REST and GraphQL modes

apiMode: "auto" is recommended:

  • With a githubToken, the Actor uses GitHub GraphQL.
  • Without a token, the Actor uses GitHub REST.

Choose rest to force REST even when a token is present. Choose graphql to force GraphQL; this mode requires githubToken.

GraphQL queries the default branch through defaultBranchRef or the selected branch, requests only the fields used in the dataset, and follows forward cursors. REST remains available for backward-compatible tokenless tasks.

Neither route is an unlimited-rate workaround. The Actor honors GitHub retry/reset guidance, bounds retries, detects GraphQL errors returned with HTTP 200, and stops admitting work before the Apify timeout.

Repository and organization behavior

Explicit repositories retain sourceType: "repository".

Repositories discovered from organization use sourceType: "organization". If the same repository appears in both places, the explicit source wins and the row is not duplicated for discovery provenance.

maxRepositories limits only organization discovery. Explicit repository inputs remain separate, and all sources are deduplicated by canonical owner/repo.

Date and branch filters

since and until accept ISO dates or timestamps:

{
"repositories": ["apify/crawlee"],
"since": "2026-01-01T00:00:00Z",
"until": "2026-06-30T23:59:59Z"
}

Use branch for a non-default branch:

{
"repositories": ["owner/repository"],
"branch": "release/v1",
"maxCommitsPerRepo": 100
}

If a valid date range or empty repository has no matching commits, the run succeeds with zero rows and records the reason in RUN_SUMMARY. If every repository is unavailable or fails, the run fails instead of presenting a misleading empty success.

Noreply filtering and email deduplication

GitHub users can author commits with privacy-protecting addresses ending in users.noreply.github.com.

  • Keep includeNoreplyEmails: false to exclude these addresses.
  • Set it to true for complete public commit evidence.

With dedupeByEmail: true, only the first row for each case-insensitive email is saved across all repositories. Disable it when you need a row for every inspected commit that contains an included email.

Output fields

FieldTypeDescription
repositorystringCanonical owner/repo.
ownerstringRepository owner login.
repostringRepository name.
commitShastringGit commit object ID.
commitUrlstringPublic source commit URL.
authorNamestring or nullAuthor name stored in the commit.
authorEmailstringAuthor email stored in the public commit.
isNoreplyEmailbooleanWhether the address is a GitHub noreply address.
githubLoginstring or nullLinked GitHub login when available.
profileUrlstring or nullLinked public GitHub profile URL.
committedAtstring or nullCommit author timestamp.
messageSnippetstring or nullFirst 240 normalized characters of the commit message.
sourceTypestringrepository or organization.
scrapedAtstringRow-save timestamp.
apiModestringrest or graphql.

Run summary and pending work

The default key-value store contains:

  • RUN_SUMMARY — route, requested/completed/failed repository counts, processed commits, saved rows, skip counts, empty reason, warnings, and classified repository errors.
  • PENDING_WORK — the repository and page/cursor checkpoint used when work remains.

Rows are saved progressively. Saving a paid row and charging the item event are one operation. If a later repository fails or the active-work deadline is reached, completed rows remain available.

On a successful complete run, PENDING_WORK is cleared. On a forced cutoff, the run fails honestly with partial output and a resumable checkpoint.

Pricing

This Actor uses pay-per-event pricing. Prices below match the active Actor configuration.

EventCharged unitPrice
startOnce per run after valid input$0.005
EventFree / no discountStarter / BronzeScale / SilverBusiness / GoldCustom / PlatinumCustom / Diamond
item per saved row$0.000030184$0.000026247$0.000020472$0.000015748$0.000010499$0.000007349

Only successfully saved contributor-email rows receive the item charge. Repository errors, malformed records, filtered commits, duplicate rows, and summary/checkpoint records are not charged as items.

Apify can separately charge platform usage for compute, storage, and data transfer. This Actor does not use a paid proxy by default.

Who is this for?

This Actor is designed for teams that have authorization to work with public Git commit metadata, including:

  • developer relations and open-source program teams;
  • recruiting and talent research teams;
  • repository maintainers and engineering analytics teams;
  • security, compliance, and software-supply-chain analysts;
  • data teams building approved contributor directories or audits.

It is not a private-email discovery service. It exports only author identity fields already attached to public commits returned by GitHub's official APIs.

Common workflows

Authorized contributor research: Export recent public commit evidence for repositories you are permitted to analyze.

Developer relations: Identify contributors across owned or authorized open-source projects and retain a source commit for review.

Repository auditing: Include noreply addresses and disable deduplication to examine commit-level author metadata.

Organization reporting: Discover a bounded set of public organization repositories and export a unified dataset.

Recurring monitoring: Use since with Apify schedules to collect recent commit-author records at a controlled cadence.

Schedule recurring exports

Save a task with a bounded repository list and date filter, then attach it to an Apify schedule.

For recurring jobs:

  • Keep repository and commit limits conservative.
  • Use a GitHub token for GraphQL or larger workloads.
  • Advance since in your automation when you need incremental windows.
  • Check RUN_SUMMARY before treating a partial multi-repository result as complete.
  • Use the dataset's stable repository/SHA/email combination when merging runs.

API usage

Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('fetch_cat/github-contributor-email-scraper').call({
repositories: ['apify/crawlee'],
maxCommitsPerRepo: 100,
includeNoreplyEmails: false,
dedupeByEmail: true,
apiMode: 'auto',
githubToken: process.env.GITHUB_TOKEN
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

Python

import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("fetch_cat/github-contributor-email-scraper").call(run_input={
"repositories": ["apify/crawlee"],
"maxCommitsPerRepo": 100,
"includeNoreplyEmails": False,
"dedupeByEmail": True,
"apiMode": "rest",
})
items = client.dataset(run["defaultDatasetId"]).list_items().items
print(items)

cURL

curl -X POST \
"https://api.apify.com/v2/acts/fetch_cat~github-contributor-email-scraper/runs" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"repositories": ["apify/crawlee"],
"maxCommitsPerRepo": 100,
"apiMode": "rest"
}'

Pass Apify authentication in the header so it does not enter URL history or request logs. Put githubToken in the secret input field rather than source code.

Use with MCP and AI agents

This Actor works through the official Apify MCP server. The focused endpoint exposes only this Actor:

https://mcp.apify.com?tools=fetch_cat/github-contributor-email-scraper

Claude Code:

claude mcp add apify-github-contributor-email \
--url "https://mcp.apify.com?tools=fetch_cat/github-contributor-email-scraper"

Claude Desktop:

{
"mcpServers": {
"apify-github-contributor-email": {
"url": "https://mcp.apify.com?tools=fetch_cat/github-contributor-email-scraper"
}
}
}

Example prompts:

  • "Run the GitHub Contributor Email Scraper for apify/crawlee with 50 commits and exclude noreply addresses."
  • "Export commit-level author evidence from these authorized repositories without email deduplication."
  • "Scan the default branches of these repositories since 2026-01-01 and summarize RUN_SUMMARY before using the rows."

Tips for better results

  • Start with one repository and 20-100 commits.
  • Use apiMode: "auto" with a GitHub token for GraphQL.
  • Use branch only when you know the requested branch exists across the repositories.
  • Keep dedupeByEmail enabled for one row per address.
  • Disable deduplication for commit-level evidence and expect more rows.
  • Review RUN_SUMMARY for partial repository failures and rate-limit diagnostics.
  • If GitHub returns a reset time, wait rather than immediately launching repeated large runs.

Limits and practical notes

  • Public repositories only.
  • The selected branch history is linear and follows GitHub's API semantics.
  • A commit email may not belong to a deliverable mailbox.
  • GitHub may return no linked login even when the commit has an author name and email.
  • Noreply addresses intentionally protect a contributor's contact address.
  • Rewritten Git history can change or remove earlier commit evidence.
  • REST and GraphQL have separate primary limits but share secondary protections.
  • The Actor retries transient failures only within a bounded run deadline.
  • It does not use residential proxies, browser automation, cookies, or private-repository permissions.

Legality and responsible use

Use this Actor only within the GitHub authorization and legal basis that applies to your workflow. Follow applicable privacy, employment, communications, anti-spam, security, and data-retention requirements.

Do not assume a public commit email is consent for unsolicited messaging. Maintain appropriate suppression, removal, and do-not-contact handling for downstream systems.

FAQ

Does it require a GitHub token?

No for small public REST runs. GraphQL requires githubToken, and a token also provides higher authenticated API limits.

Does it scrape private GitHub emails?

No. It exports the author email stored in a public Git commit. It does not infer hidden addresses or bypass GitHub privacy controls.

Can it export CSV, Excel, or JSON?

Yes. Use the Apify dataset export controls or API to download CSV, Excel, JSON, XML, and other supported formats.

Can it scan an entire organization?

It can discover up to 100 public organization repositories per run. The default is 50. Combine organization with explicit repositories when needed.

Why are some profile fields null?

GitHub does not always link a Git commit author to a GitHub account. The commit email, author name, and linked account are separate pieces of metadata.

Why did a valid run return zero rows?

The date range may contain no commits, the repository may be empty, every matching commit may lack an email, or all addresses may have been excluded as noreply. RUN_SUMMARY.emptyReason explains the observed case.

What happens when one repository fails?

Completed repositories remain saved. The failed repository is classified in RUN_SUMMARY.repositoryErrors. The whole run fails only when every repository fails or when the shared run deadline stops unfinished work.

Is GraphQL always faster?

Not necessarily. It provides a compact authenticated query and a separate primary quota, but query cost, node limits, timeouts, and secondary limits still apply. auto chooses GraphQL only when a token is available.

Support

If a run fails or the output looks wrong, open an issue or report a bug from the Actor page.

Please include:

  • Apify run ID or run URL
  • Input JSON with secret tokens removed
  • Expected output
  • Actual output
  • One reproducible public repository URL

This lets support distinguish GitHub availability, rate limits, branch/date filters, schema changes, and Actor behavior without exposing credentials.

For one reproducible public URL, use https://github.com/apify/crawlee and compare against the Crawlee contributor-email example.