GitLab Email Scraper
Pricing
from $2.49 / 1,000 results
GitLab Email Scraper
GitLab Email Scraper SD - GitLab Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from GitLab results by keyword, location and email domain - GitLab email extractor.
Pricing
from $2.49 / 1,000 results
Rating
0.0
(0)
Developer
Neuro Scraper
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
14 days ago
Last modified
Categories
Share
GitLab Email Scraper - find public developer and maintainer emails on GitLab
The GitLab Email Scraper is an Apify Actor that collects publicly indexed contact emails tied to GitLab users, groups, projects and GitLab Pages sites. Give it keywords, an optional location and your email domains, and it returns a structured lead dataset.
GitLab is where a lot of engineering work lives in the open: project README files, CONTRIBUTING and SECURITY docs, CI configuration notes, merge request discussions, group landing pages and documentation published to gitlab.io.
Wherever a maintainer has written a contact address on one of those public pages and Google has indexed it, the GitLab Email Scraper can pick it up and hand it to you as a row with a handle and a profile URL.
What the GitLab Email Scraper reads - and what it never touches
The GitLab Email Scraper reads only Google's public index of gitlab.com and gitlab.io. It builds site: queries, fetches Google result pages through the Apify GOOGLE_SERP proxy, and extracts emails from the result titles and snippets.
It does not read git history, clone repositories, or call the GitLab API - which is important, because the GitLab Users API only exposes other people's email addresses to instance administrators.
There is no login, no access token, no browser, no JavaScript rendering and no cookies. Self-managed GitLab instances on private domains are out of scope; only gitlab.com and gitlab.io are queried.
Who uses the GitLab Email Scraper
Technical recruiters who want engineers with demonstrable DevOps and CI experience, developer relations teams mapping an ecosystem, agencies doing developer lead generation, and maintainers looking for peer projects to partner with.
If you have ever searched for a way to find a GitLab user's email address and hit the admin-only API wall, this Actor is the practical alternative that works from public data.
Key features of the GitLab Email Scraper
| Feature | What it does |
|---|---|
Google site: dorking | Queries gitlab.com and gitlab.io through the Apify GOOGLE_SERP proxy |
| Query expansion | Base, quoted and intitle: variants plus one variant per query modifier; base queries run first |
| Domain-filtered extraction | Keeps only emails ending in your customDomains values |
| Global deduplication | One row per unique email address across the whole run |
| Obfuscation handling | Understands name [at] domain [dot] com, name (at) domain, name @ domain.com, domain .com, zero-width characters and the full-width @ |
| Junk filter | Rejects placeholders such as email@, yourname@, test@, xxx@ and single-character locals |
| Boundary-correct matching | @gmail.com never matches inside @gmail.company or @gmail.com.br |
| Soft-wrap repair | Discards a hit that is only the tail of another email in the same block |
| Structural parsing | Finds the <h3> title then the smallest surrounding block, not Google's CSS class names |
| Whole-page fallback | A Google layout change degrades the run to "emails without account details", not "no emails" |
| Block detection | CAPTCHA, "unusual traffic" and consent pages are detected and retried rather than counted as empty |
| Retries and backoff | Up to 3 attempts per page, exponential backoff, a fresh proxy session per request |
| Requeue of failures | Blocked or failed queries are re-queued once at the end of the run |
| Async concurrency | An asyncio worker pool with a shared stop signal on maxEmails |
| Resumable state | Key-value-store progress keyed by an input hash, with throttled saves plus saves on PERSIST_STATE, MIGRATING and ABORTING |
| Streaming output | Leads are pushed to the dataset as they are found, so an aborted run keeps what it collected |
| Run summary | Logs pages fetched, blocked pages, retries and emails per page |
How the GitLab Email Scraper works
The GitLab Email Scraper pipeline is six steps long, with no browser and no authentication anywhere in it.
- The GitLab Email Scraper reads your input: keywords, location, email domains and limits.
- It builds Google queries with the
site:operator, for examplesite:gitlab.com maintainer "@gmail.com" "Amsterdam". - It fetches Google result pages asynchronously with
aiohttpthrough the Apify GOOGLE_SERP proxy. - It parses each result block structurally, locating the
<h3>title and the smallest block around it. - It extracts email addresses from the block text using a domain-filtered regex.
- It deduplicates globally and pushes every lead straight into the Apify dataset.
Why query expansion matters here
Google caps a single query at roughly 300 results, so one query has a hard ceiling regardless of how many pages you allow.
The GitLab Email Scraper works around that by combining every keyword with every email domain in several phrasings, giving each variant its own result budget.
The default queryModifiers - email, contact, maintainer, author, support - are the words that genuinely appear next to addresses in GitLab project documentation and group descriptions.
What the GitLab Email Scraper does not do
No login, no personal access token, no GitLab API, no repository cloning, no commit-history parsing, no JavaScript rendering. It is an independent Apify Actor and is not affiliated with or endorsed by GitLab.
GitLab Email Scraper input fields
keywords is the only required field in the GitLab Email Scraper. Every default below is the one shipped in the Actor's input schema.
| Field | Type | Default | Meaning |
|---|---|---|---|
keywords | array (required) | ["developer", "maintainer"] | Search terms describing the GitLab accounts you want (niche, job title, industry) |
location | string | "" | Optional location phrase added to every query |
customDomains | array | ["@gmail.com", "@yahoo.com"] | Only emails on these domains are kept; the leading @ is optional |
maxEmails | integer 1-10000 | 20 | Stop after this many unique emails |
countryCode | string | "" | Two-letter country for the search proxy (US, GB, DE...) |
expandQueries | boolean | true | Search each keyword x domain pair in several phrasings |
queryModifiers | array | ["email", "contact", "maintainer", "author", "support"] | Extra words combined with each keyword when expansion is on |
maxPagesPerQuery | integer 1-50 | 30 | Page cap per query |
maxConcurrency | integer 1-20 | 5 | Parallel queries |
Example GitLab Email Scraper input
{"keywords": ["ci cd engineer", "kubernetes maintainer", "platform engineer"],"location": "Amsterdam","customDomains": ["@gmail.com", "@protonmail.com"],"maxEmails": 400,"countryCode": "NL","expandQueries": true,"queryModifiers": ["email", "contact", "maintainer", "author", "support"],"maxPagesPerQuery": 30,"maxConcurrency": 5}
Tuning the GitLab Email Scraper
Narrow keywords out-perform broad ones in the GitLab Email Scraper. gitlab runner maintainer beats developer, because it matches the vocabulary people actually use on project pages.
When a run comes back thin, add domains before you add keywords. Each additional domain multiplies the number of distinct queries and therefore the total result budget.
Keep expandQueries on unless you are running a quick smoke test - turning it off collapses the GitLab Email Scraper to a single query per keyword and domain pair.
GitLab Email Scraper output fields
Every dataset item the GitLab Email Scraper produces has all 14 fields. Nothing is omitted; values are empty or null when Google did not expose them.
| Field | Meaning |
|---|---|
network | Platform name |
keyword | The keyword that produced the lead |
query | The exact Google query used |
title | Raw result title |
accountName | Account label Google prints (handle, display name, or group name) |
fullName | Display name parsed from a profile-style title; empty for project or doc pages |
username | URL-safe GitLab handle or namespace when one is exposed; otherwise null |
profileUrl | Canonical https://gitlab.com/{username} URL when a handle is known; otherwise empty |
url | Direct platform link when exposed, else the profile URL |
description | Bio or snippet text, cleaned of labels and engagement counters |
email | Lower-cased email address |
emailDomain | The matched domain, for example @gmail.com |
possiblyTruncated | true when Google's snippet ellipsis touched the email - verify before sending |
foundAt | ISO 8601 UTC timestamp |
Example GitLab Email Scraper output
[{"network": "GitLab","keyword": "ci cd engineer","query": "site:gitlab.com ci cd engineer \"@gmail.com\" \"Amsterdam\"","title": "Daan Verhoeven (dverhoeven) - GitLab","accountName": "dverhoeven","fullName": "Daan Verhoeven","username": "dverhoeven","profileUrl": "https://gitlab.com/dverhoeven","url": "https://gitlab.com/dverhoeven","description": "Platform engineer in Amsterdam. CI pipelines, runners, Terraform. Contact daan.verhoeven.dev@gmail.com","email": "daan.verhoeven.dev@gmail.com","emailDomain": "@gmail.com","possiblyTruncated": false,"foundAt": "2026-08-31T10:02:41.117Z"},{"network": "GitLab","keyword": "kubernetes maintainer","query": "site:gitlab.com intitle:\"kubernetes maintainer\" \"@gmail.com\"","title": "helios-infra / helios-operator - CONTRIBUTING","accountName": "helios-infra","fullName": "","username": "helios-infra","profileUrl": "https://gitlab.com/helios-infra","url": "https://gitlab.com/helios-infra/helios-operator","description": "Maintainers reachable at helios.maintainers@gmail.com for security reports and release ...","email": "helios.maintainers@gmail.com","emailDomain": "@gmail.com","possiblyTruncated": true,"foundAt": "2026-08-31T10:02:55.640Z"},{"network": "GitLab","keyword": "platform engineer","query": "site:gitlab.io platform engineer contact \"@protonmail.com\"","title": "Handbook - Contact","accountName": "Handbook","fullName": "","username": null,"profileUrl": "","url": "https://northwind-platform.gitlab.io/handbook/contact/","description": "Questions about the platform team? Write to northwind.platform@protonmail.com","email": "northwind.platform@protonmail.com","emailDomain": "@protonmail.com","possiblyTruncated": false,"foundAt": "2026-08-31T10:03:12.008Z"}]
The third row shows the honest edge case: a GitLab Pages site on gitlab.io publishes a real address but no GitLab handle, so username is null and profileUrl is empty. The GitLab Email Scraper reports that rather than guessing a URL.
Use cases for the GitLab Email Scraper
| Use case | How the GitLab Email Scraper helps |
|---|---|
| Technical recruiting | Source engineers by stack, role and city, with a profile URL you can review first |
| DevOps and platform sourcing | GitLab skews toward CI/CD and infrastructure work, so keyword targeting is unusually precise |
| Developer relations | Find maintainers of adjacent projects before a launch, beta or integration push |
| Open-source outreach | Reach the contact addresses published in README, CONTRIBUTING and SECURITY docs |
| Developer lead generation | Fill a devtools or API pipeline with accounts matching a technical niche |
| Partnership research | Identify groups and organisations shipping tooling in your category |
| Security disclosure | Collect published security contacts for projects in your dependency tree |
| Ecosystem research | Measure how many projects in a niche publish a public contact channel at all |
Sourcing infrastructure engineers
GitLab's centre of gravity is CI/CD, runners and self-hosted delivery, so the GitLab Email Scraper is a sharper instrument than a general web scrape when you are hiring for platform roles.
Pair it with the GitHub Email Scraper to cover engineers who keep their public code on one host and their day job on the other.
Building a launch list
DevRel teams commonly run the GitLab Email Scraper alongside the Docker Hub Email Scraper, since the people publishing pipelines and the people publishing images are largely the same population.
Expected results from the GitLab Email Scraper
In live test runs against Google, about 8 out of 10 parsed GitLab results carried an account identity - a handle, a display name, or both - and profile URLs resolved reliably.
That is the strongest coverage in this Developer & Technology family, because GitLab's user and namespace titles are consistently formatted in Google's index.
GitLab Email Scraper rows without a handle still carry email, emailDomain, url, description and query. Total volume always depends on your keywords, domains and location; no yield is guaranteed.
Limitations of the GitLab Email Scraper
- Publicly indexed emails only. An address is findable only when it is already visible in Google's index. Private projects, unindexed pages and addresses users never published are unreachable.
- No git history, no GitLab API. Commit author emails and API responses are not read. Only Google result titles and snippets are parsed.
- Only
gitlab.comandgitlab.io. Self-managed GitLab instances on private domains are not queried. - Google's ~300-result cap. One query returns roughly 300 results at most, which is exactly why
expandQueriesexists. possiblyTruncated. Whentrue, Google's snippet ellipsis touched the email and it may be cut off. Verify those rows before sending.usernameandprofileUrlcan be empty. They are filled only when GitLab exposes a handle in the result. Somegitlab.iopages and documentation files show only a display name, leavingaccountNamepopulated butusernamenull.- Apify GOOGLE_SERP proxy required. The GitLab Email Scraper cannot run without Apify proxy credentials.
- Free plan cap. Free Apify plans are limited to 100 emails per run; paid plans are uncapped.
- Variable results. Google's index moves, so two runs a month apart will not be identical.
Responsible use
Emails the GitLab Email Scraper returns were published on GitLab project pages for project correspondence: bug reports, security disclosures, release questions, contribution coordination.
If you contact people in the EU or UK, GDPR applies. Have a lawful basis, identify yourself clearly, say where you found the address, and honour opt-out requests straight away.
Respect the norms of the platform and of individual projects. A README that says "no recruiters" means no recruiters, and relevance beats volume every single time.
GitLab Email Scraper FAQ
Does the GitLab Email Scraper use the GitLab API?
No. It does not call the GitLab API, read git history or clone repositories. It parses Google search results for already-indexed public pages on gitlab.com and gitlab.io.
Do I need a GitLab account or personal access token?
No. There is no authentication of any kind. The only credential involved is your Apify account's GOOGLE_SERP proxy access.
Can it search a self-managed GitLab instance?
No. The GitLab Email Scraper queries gitlab.com and gitlab.io only. A private instance on your own domain is not covered.
Where do the emails come from?
The GitLab Email Scraper takes them from the text Google prints in titles and snippets: profile bios, group and project descriptions, README and CONTRIBUTING files, merge request and issue text, and GitLab Pages documentation.
How many emails can one GitLab Email Scraper run return?
maxEmails accepts 1 to 10000 and defaults to 20. Free Apify plans are capped at 100 emails per run; paid plans are uncapped.
Why is username null on some rows?
Because Google did not expose a GitLab handle for that result. The row is still a usable lead with email, accountName, url and description.
What does possiblyTruncated: true mean?
Google's snippet ellipsis touched the address, so it may be incomplete. Check those rows before you use them.
Can the GitLab Email Scraper collect only company-domain emails?
Yes. Set customDomains to the domains you want, for example ["@acme.io", "@example.com"]. The leading @ is optional and only matching addresses are kept.
Does it handle name [at] gmail [dot] com style obfuscation?
Yes. The extractor normalises [at], (at), spaced @, spaced .com, zero-width characters and the full-width @, and filters placeholders such as test@ and yourname@.
How do I target one country?
Set countryCode to a two-letter code such as US, GB or DE, and put a city or region into location so it is appended to every query.
What happens if a GitLab Email Scraper run is interrupted or migrated?
State is saved in the key-value store keyed by a hash of your input, including on PERSIST_STATE, MIGRATING and ABORTING. Leads already pushed to the dataset are never lost.
Will a Google redesign break it?
Parsing is structural rather than class-name based and there is a whole-page fallback parser, so a layout change degrades the GitLab Email Scraper to "emails without account details" instead of nothing.
Related Actors
The GitLab Email Scraper belongs to a Developer & Technology family of Apify Actors that apply the same method to different platforms. Running several gives you ecosystem coverage instead of a single-site snapshot.
| Actor | What it collects |
|---|---|
| GitLab Email and Phone Number Scraper | Emails and phone numbers from GitLab |
| GitLab Phone Number Scraper | Public phone numbers from GitLab |
| App Store Email Scraper | Public contact emails from App Store |
| Atlassian Marketplace Email Scraper | Public contact emails from Atlassian Marketplace |
| Bitbucket Email Scraper | Public contact emails from Bitbucket |
| Chrome Web Store Email Scraper | Public contact emails from Chrome Web Store |
| CodePen Email Scraper | Public contact emails from CodePen |
| Confluence Email Scraper | Public contact emails from Confluence |
| Dev.to Email Scraper | Public contact emails from DEV Community |
| Docker Hub Email Scraper | Public contact emails from Docker Hub |
| Figma Community Email Scraper | Public contact emails from Figma Community |
| Firefox Add-ons Email Scraper | Public contact emails from Firefox Add-ons |
| GitHub Email Scraper | Public contact emails from GitHub |
| Google Play Email Scraper | Public contact emails from Google Play |
| Hashnode Email Scraper | Public contact emails from Hashnode |
| HubSpot Marketplace Email Scraper | Public contact emails from HubSpot Marketplace |
| Hugging Face Email Scraper | Public contact emails from Hugging Face |
| Jira Email Scraper | Public contact emails from Jira |
| Maven Central Email Scraper | Public contact emails from Maven Central |
| Microsoft AppSource Email Scraper | Public contact emails from Microsoft AppSource |
| Salesforce AppExchange Email Scraper | Public contact emails from Salesforce AppExchange |
| Shopify App Store Email Scraper | Public contact emails from Shopify App Store |
| Slack App Directory Email Scraper | Public contact emails from Slack App Directory |
| SourceForge Email Scraper | Public contact emails from SourceForge |
| Stack Overflow Email Scraper | Public contact emails from Stack Overflow |
| Unity Asset Store Email Scraper | Public contact emails from Unity Asset Store |
| Unreal Engine Marketplace Email Scraper | Public contact emails from Unreal Engine Marketplace |
| WordPress Plugin Directory Email Scraper | Public contact emails from WordPress Plugin Directory |
| WordPress Theme Directory Email Scraper | Public contact emails from WordPress Theme Directory |
| Zapier App Directory Email Scraper | Public contact emails from Zapier App Directory |
| App Store Email and Phone Number Scraper | Emails and phone numbers from App Store |
| Atlassian Marketplace Email and Phone Number Scraper | Emails and phone numbers from Atlassian Marketplace |
| Bitbucket Email and Phone Number Scraper | Emails and phone numbers from Bitbucket |
| Chrome Web Store Email and Phone Number Scraper | Emails and phone numbers from Chrome Web Store |
| CodePen Email and Phone Number Scraper | Emails and phone numbers from CodePen |
| Confluence Email and Phone Number Scraper | Emails and phone numbers from Confluence |
| DEV Community Email and Phone Number Scraper | Emails and phone numbers from DEV Community |
| Docker Hub Email and Phone Number Scraper | Emails and phone numbers from Docker Hub |
| Figma Community Email and Phone Number Scraper | Emails and phone numbers from Figma Community |
| Firefox Add-ons Email and Phone Number Scraper | Emails and phone numbers from Firefox Add-ons |
| GitHub Email and Phone Number Scraper | Emails and phone numbers from GitHub |
| Google Play Email and Phone Number Scraper | Emails and phone numbers from Google Play |
Leave a review
If the GitLab Email Scraper saved you time, please leave a star rating and a short review on the Actor page.
Reviews are how other buyers judge whether a tool works, and they tell us which features to build next.
If something did not work, email neurodata.apify@gmail.com instead - bugs get fixed faster than they get complained about.
Support
Need help, found a bug, or want a custom build of the GitLab Email Scraper tuned to your own domain list? Email neurodata.apify@gmail.com.