Maven Central Email Scraper
Pricing
from $2.49 / 1,000 results
Maven Central Email Scraper
Maven Central Email Scraper SD - Maven Central Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from Maven Central results by keyword, location and email domain - Maven Central email extractor.
Pricing
from $2.49 / 1,000 results
Rating
0.0
(0)
Developer
Neuro Scraper
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
7 days ago
Last modified
Categories
Share
Maven Central Email Scraper — project contacts from indexed artifact listings
Maven Central Email Scraper collects publicly indexed contact emails that appear alongside Java and JVM artifact listings in Google's search index. You supply keywords, the email domains you care about and an optional location, and it returns a deduplicated dataset of project contacts.
Read this before you run it: Maven listings identify artifacts and organisations rather than people, so rows usually carry a project contact instead of a personal handle. In measured runs the Maven Central Email Scraper resolved account identity on only 1 row out of 10, and that is the honest expectation to plan around.
That is not a defect in the Actor — it is how the JVM ecosystem works. A published artifact is identified by a groupId and artifactId, and the contact printed next to it is very often a mailing list, an organisation address or a project inbox rather than an individual developer's.
Unlike most Actors in this family, the Maven Central Email Scraper searches three domains rather than one: mvnrepository.com, central.sonatype.com and search.maven.org. Those three sites index the same repository from different angles, and each renders project metadata a little differently.
The Maven Central Email Scraper does not use any Maven or Sonatype API, does not download JARs and does not read pom.xml or any other build file. Everything it returns comes from Google result blocks — titles, snippets and site labels — fetched through the Apify GOOGLE_SERP proxy.
Who the Maven Central Email Scraper is for
Java tooling vendors, security researchers doing coordinated disclosure on JVM libraries, supply-chain analysts, enterprise open-source programme offices and consultancies selling into JVM shops all want the same thing: a way to reach the organisation behind a published artifact.
If you also work in other package ecosystems, the Maven Central Email Scraper pairs directly with the npm Email Scraper and the PyPI Email Scraper, both of which yield personal maintainer contacts far more often.
Key features of the Maven Central Email Scraper
Everything the Actor can do is in this table. If a capability is not listed, the Maven Central Email Scraper does not have it.
| Feature | What it does |
|---|---|
Three-domain site: targeting | Queries are scoped to mvnrepository.com, central.sonatype.com and search.maven.org |
| Artifact URL building | Constructs profileUrl as https://mvnrepository.com/artifact/{username} when a coordinate-style handle is exposed |
| Query expansion | Runs each keyword as a base query, a quoted phrase, an intitle: query and one variant per query modifier |
| Domain filtering | Keeps only emails ending in the domains you list, with boundary-correct matching |
| Obfuscation handling | Understands name [at] domain [dot] com, name (at) domain, name @ domain.com, domain .com, zero-width characters and the full-width @ |
| Junk filter | Rejects placeholders such as email@, yourname@, test@, xxx@ and single-character locals |
| Soft-wrap repair | Drops a hit that is only the tail of another email in the same result block |
| Global deduplication | One row per unique email address across all three domains, every query and every page |
| Truncation flag | Sets possiblyTruncated: true when Google's snippet ellipsis touched the address |
| Structural parsing | Locates the <h3> title then the smallest surrounding block, rather than depending on Google's CSS class names |
| Whole-page fallback | A markup change degrades the run to "emails without account details" instead of "no emails" |
| Block detection | CAPTCHA, "unusual traffic" and consent pages are detected and retried, not counted as empty |
| Retries and backoff | Up to 3 attempts per page with exponential backoff and a fresh proxy session per request |
| Requeue of failures | Blocked or failed queries are re-queued once at the end of the run |
| Resumable state | Progress lives in the key-value store keyed by a hash of your input, saved on PERSIST_STATE, MIGRATING and ABORTING |
| Concurrency control | An asyncio worker pool with a shared stop signal on maxEmails |
| Streaming output | Every lead is pushed to the dataset the moment it is found |
How the Maven Central Email Scraper works
Six steps, all of them visible in the run log.
- The Maven Central Email Scraper reads your input: keywords, location, email domains and limits.
- It builds Google queries with the
site:operator against each of its three domains, for examplesite:mvnrepository.com json parser contact "@gmail.com". - It fetches Google result pages asynchronously with
aiohttpthrough the Apify GOOGLE_SERP proxy. - It parses each result block structurally — find the
<h3>, then take the smallest block containing it. - It extracts email addresses from that block's text with a domain-filtered regular expression.
- It deduplicates globally and pushes each new lead straight into the Apify dataset.
There is no browser, no JavaScript rendering, no authentication and no cookies. The Maven Central Email Scraper only ever sees what Google already shows the public.
Query expansion matters more here than almost anywhere else. Google caps a single query at roughly 300 results, and with contact addresses appearing sparsely across artifact listings, the extra phrasings are often the difference between an empty dataset and a usable one.
Base queries run first, so the strongest matches land before the long tail. Because three domains are in play, expect the Maven Central Email Scraper to spend a noticeable share of a run on the two Sonatype-operated sites.
Input fields
Every input field of the Maven Central Email Scraper is listed below exactly as defined in the Actor's input schema. Only keywords is required.
| Field | Type | Default | Meaning |
|---|---|---|---|
keywords | array (required) | ["java library", "artifact"] | Search terms describing the Maven Central listings you want |
location | string | "" | Optional location phrase added to every query |
customDomains | array | ["@gmail.com", "@yahoo.com"] | Only emails on these domains are kept; the @ is optional |
maxEmails | integer 1–10000 | 20 | Stop after this many unique emails |
countryCode | string | "" | Two-letter country for the search proxy (US, GB, DE…) |
expandQueries | boolean | true | Search each keyword × domain pair in several phrasings |
queryModifiers | array | ["email", "contact", "maintainer", "author", "support"] | Extra words combined with each keyword when expansion is on |
maxPagesPerQuery | integer 1–50 | 30 | Page cap per query |
maxConcurrency | integer 1–20 | 5 | Parallel queries |
The queryModifiers default is tuned for Maven Central. contact, maintainer and author are the words that appear near a published address on an artifact page, so they carry most of the weight in a run.
Example input for the Maven Central Email Scraper
{"keywords": ["json parser", "spring boot starter", "kotlin coroutines library"],"location": "","customDomains": ["@gmail.com", "@yahoo.com"],"maxEmails": 300,"countryCode": "US","expandQueries": true,"queryModifiers": ["email", "contact", "maintainer", "author", "support"],"maxPagesPerQuery": 30,"maxConcurrency": 5}
Describe the artifact, not a person. logging framework, jdbc driver, gradle plugin and protobuf runtime match how artifact pages present themselves far better than any job title.
Output fields
Every dataset item produced by the Maven Central Email Scraper carries all fourteen fields below. Nothing is omitted; fields with no value are empty or null.
| Field | Meaning |
|---|---|
network | Platform name |
keyword | The keyword that produced the lead |
query | The exact Google query used |
title | Raw result title |
accountName | Account label Google prints — on Maven Central this is normally an artifact or organisation |
fullName | Display name parsed from a profile-style title; empty for other result types |
username | A coordinate-style handle when Google exposes one; otherwise null — frequently null here |
profileUrl | https://mvnrepository.com/artifact/{username} when a handle is known; otherwise empty |
url | Direct listing link when exposed, else the profile URL |
description | Artifact snippet, cleaned of labels and counters |
email | Lower-cased email address |
emailDomain | The matched domain, e.g. @gmail.com |
possiblyTruncated | true when Google's snippet ellipsis touched the email — verify before sending |
foundAt | ISO 8601 UTC timestamp |
Example output from the Maven Central Email Scraper
The first row below is the good case. The second is the common case: a real, usable project address with no handle attached.
[{"network": "Maven Central","keyword": "json parser","query": "site:mvnrepository.com json parser contact \"@gmail.com\"","title": "Maven Repository: io.harborline » harborline-json","accountName": "io.harborline","fullName": "","username": "io.harborline","profileUrl": "https://mvnrepository.com/artifact/io.harborline","url": "https://mvnrepository.com/artifact/io.harborline/harborline-json","description": "Streaming JSON parser for the JVM. Project contact: harborline.oss@gmail.com","email": "harborline.oss@gmail.com","emailDomain": "@gmail.com","possiblyTruncated": false,"foundAt": "2026-08-31T13:11:26Z"},{"network": "Maven Central","keyword": "spring boot starter","query": "site:central.sonatype.com \"spring boot starter\" maintainer \"@yahoo.com\"","title": "Maven Central: northgate-metrics-starter","accountName": "northgate-metrics-starter","fullName": "","username": null,"profileUrl": "","url": "https://central.sonatype.com/artifact/com.northgate/northgate-metrics-starter","description": "Metrics auto-configuration starter. Developer mailing list: northgate.dev.list@yahoo.com","email": "northgate.dev.list@yahoo.com","emailDomain": "@yahoo.com","possiblyTruncated": true,"foundAt": "2026-08-31T13:14:02Z"}]
Rows like the second one are the norm rather than the exception on this platform. They are still useful: email, accountName, url and description are all populated by the Maven Central Email Scraper, you simply do not get a handle.
Use cases for the Maven Central Email Scraper
| Use case | How the Maven Central Email Scraper helps |
|---|---|
| Coordinated vulnerability disclosure | Reach the project inbox published next to a JVM artifact |
| Java tooling and platform sales | Target organisations publishing artifacts in the space your product serves |
| Supply-chain and dependency research | Map which organisations publish which corner of the repository |
| Enterprise open-source programme outreach | Contact the teams behind libraries your organisation depends on |
| Library adoption or handover | Reach the owning organisation about an artifact you would like to maintain |
| Sponsorship and funding programmes | Identify JVM projects with a published contact channel |
| Licensing and compliance queries | Find the correspondence address for an artifact's owning organisation |
| CRM enrichment | Attach a published project contact to an organisation you already track |
Because most JVM libraries are developed on a public forge, the Maven Central Email Scraper works well next to the GitHub Email Scraper, which is far more likely to give you a named individual.
Tips to get more from the Maven Central Email Scraper
Feed the Maven Central Email Scraper artifact functions, not people. connection pool, bytecode manipulation, test containers, xml binding and metrics exporter all appear in indexed artifact titles; job titles essentially never do.
Widen customDomains aggressively. Because JVM projects usually publish an organisational or mailing-list address, the default @gmail.com and @yahoo.com pair is often the tightest constraint on a Maven Central Email Scraper run — add the project and company domains you actually care about.
Never turn expandQueries off here. Contact addresses are sparse across artifact listings, so the modifier variants are doing most of the work in a Maven Central Email Scraper run.
Expect to iterate. Several narrow Maven Central Email Scraper runs with different artifact vocabularies will out-perform one broad run, and because deduplication is per run you can merge the datasets afterwards.
Do not expect username on most rows. Plan your downstream pipeline around email, accountName, url and description, and treat a populated handle as a bonus.
Limitations you should know before running the Maven Central Email Scraper
Maven listings identify artifacts and organisations rather than people, so rows usually carry a project contact instead of a personal handle. Measured account-identity resolution was 1 out of 10 rows. This is the dominant limitation and it will not change with better keywords.
The Maven Central Email Scraper only finds addresses that are publicly visible in Google's index. An address that exists only inside a build file will never appear.
Google caps a single query at roughly 300 results. Query expansion mitigates that ceiling across three domains; it does not remove it.
username and profileUrl are populated only when Google's result exposes a coordinate-style handle. Most rows arrive with username: null and profileUrl: "", keeping accountName, email and url. That is a Google limitation, not a bug.
possiblyTruncated: true means Google's snippet ellipsis touched the address and it may be incomplete. Verify before sending.
The Actor requires the Apify GOOGLE_SERP proxy and cannot run without Apify proxy credentials. Free Apify plans are capped at 100 emails per run; paid plans are uncapped.
Results vary with keywords, domains, country and timing, and no volume is guaranteed. The Maven Central Email Scraper reads no repository API and no build file, so pom.xml contents are outside its reach.
Responsible use
Contact addresses published next to a Maven artifact exist so that downstream users, security researchers and packagers can correspond with the project. Use them for that kind of correspondence.
Respect GDPR and equivalent local rules, honour opt-outs immediately and follow each project's stated contact norms. Many of these are shared mailing lists, so unrelated marketing is visible to the whole project and reflects badly on you.
The Maven Central Email Scraper is a discovery tool, not a sending tool. What you do with the resulting dataset is entirely your responsibility.
Maven Central Email Scraper FAQ
Does the Maven Central Email Scraper use a repository API or read pom.xml?
No. It uses no Maven or Sonatype API, downloads no JARs and reads no build file. Every field comes from Google search result blocks fetched through the Apify GOOGLE_SERP proxy.
Which sites does it actually search?
Three: mvnrepository.com, central.sonatype.com and search.maven.org. The Maven Central Email Scraper deduplicates across all three, so a project indexed on more than one site still produces a single row per unique email.
Why is username null on most rows?
Because Maven listings identify artifacts and organisations rather than people, so Google's result often exposes no coordinate-style handle at all. Those rows still carry accountName, email, url and description.
How many rows get a resolved handle?
Roughly one in ten in measured runs. Treat a populated username and profileUrl as a bonus rather than something to build a pipeline on.
Are the emails the Maven Central Email Scraper returns personal addresses?
Usually not. Expect project inboxes, mailing lists and organisational addresses more often than an individual developer's personal email.
How is profileUrl built when a handle exists?
As https://mvnrepository.com/artifact/{username}, using the coordinate-style handle Google exposed.
Which email domains does the Maven Central Email Scraper collect?
Only the domains in customDomains, defaulting to @gmail.com and @yahoo.com. Matching is boundary-correct, so @gmail.com never matches inside @gmail.company or @gmail.com.br.
Does it handle obfuscated addresses?
Yes. The Maven Central Email Scraper normalises name [at] domain [dot] com, name (at) domain, name @ domain.com, domain .com, zero-width characters and the full-width @ before filtering by domain.
What is possiblyTruncated for?
It flags rows where Google's snippet ellipsis touched the email, so the address may be cut off. Verify those before sending.
Do I need a proxy?
Yes. The Maven Central Email Scraper runs on the Apify GOOGLE_SERP proxy and cannot operate without Apify proxy credentials.
Can I resume an interrupted run?
Yes. Progress is stored in the key-value store keyed by a hash of your input, with throttled saves that also fire on PERSIST_STATE, MIGRATING and ABORTING.
What happens if Google changes its HTML?
Parsing is structural rather than class-name based, and a whole-page fallback parser exists. A layout change degrades the Maven Central Email Scraper to "emails without account details" rather than "no emails at all".
Related Actors
The Maven Central Email Scraper is one of 32 Developer & Technology email Actors that share the same engine and differ only in the platform they target.
| Actor | What it collects |
|---|---|
| Maven Central Email and Phone Number Scraper | Emails and phone numbers from Maven Central |
| Maven Central Phone Number Scraper | Public phone numbers from Maven Central |
| App Store Email Scraper | Public contact emails from App Store |
| Atlassian Marketplace Email Scraper | Public contact emails from Atlassian Marketplace |
| Bitbucket Email Scraper | Public contact emails from Bitbucket |
| Chrome Web Store Email Scraper | Public contact emails from Chrome Web Store |
| CodePen Email Scraper | Public contact emails from CodePen |
| Confluence Email Scraper | Public contact emails from Confluence |
| Dev.to Email Scraper | Public contact emails from DEV Community |
| Docker Hub Email Scraper | Public contact emails from Docker Hub |
| Figma Community Email Scraper | Public contact emails from Figma Community |
| Firefox Add-ons Email Scraper | Public contact emails from Firefox Add-ons |
| GitHub Email Scraper | Public contact emails from GitHub |
| GitLab Email Scraper | Public contact emails from GitLab |
| Google Play Email Scraper | Public contact emails from Google Play |
| Hashnode Email Scraper | Public contact emails from Hashnode |
| HubSpot Marketplace Email Scraper | Public contact emails from HubSpot Marketplace |
| Hugging Face Email Scraper | Public contact emails from Hugging Face |
| Jira Email Scraper | Public contact emails from Jira |
| Microsoft AppSource Email Scraper | Public contact emails from Microsoft AppSource |
| Salesforce AppExchange Email Scraper | Public contact emails from Salesforce AppExchange |
| Shopify App Store Email Scraper | Public contact emails from Shopify App Store |
| Slack App Directory Email Scraper | Public contact emails from Slack App Directory |
| SourceForge Email Scraper | Public contact emails from SourceForge |
| Stack Overflow Email Scraper | Public contact emails from Stack Overflow |
| Unity Asset Store Email Scraper | Public contact emails from Unity Asset Store |
| Unreal Engine Marketplace Email Scraper | Public contact emails from Unreal Engine Marketplace |
| WordPress Plugin Directory Email Scraper | Public contact emails from WordPress Plugin Directory |
| WordPress Theme Directory Email Scraper | Public contact emails from WordPress Theme Directory |
| Zapier App Directory Email Scraper | Public contact emails from Zapier App Directory |
| App Store Email and Phone Number Scraper | Emails and phone numbers from App Store |
| Atlassian Marketplace Email and Phone Number Scraper | Emails and phone numbers from Atlassian Marketplace |
| Bitbucket Email and Phone Number Scraper | Emails and phone numbers from Bitbucket |
| Chrome Web Store Email and Phone Number Scraper | Emails and phone numbers from Chrome Web Store |
| CodePen Email and Phone Number Scraper | Emails and phone numbers from CodePen |
| Confluence Email and Phone Number Scraper | Emails and phone numbers from Confluence |
| DEV Community Email and Phone Number Scraper | Emails and phone numbers from DEV Community |
| Docker Hub Email and Phone Number Scraper | Emails and phone numbers from Docker Hub |
| Figma Community Email and Phone Number Scraper | Emails and phone numbers from Figma Community |
| Firefox Add-ons Email and Phone Number Scraper | Emails and phone numbers from Firefox Add-ons |
| GitHub Email and Phone Number Scraper | Emails and phone numbers from GitHub |
| GitLab Email and Phone Number Scraper | Emails and phone numbers from GitLab |
Leave a review
If the Maven Central Email Scraper saved you time, please leave a star rating and a short review on the Actor page.
Reviews are how other buyers judge whether a tool works, and they tell us which features to build next.
If something did not work, email neurodata.apify@gmail.com instead - bugs get fixed faster than they get complained about.
Support
Questions, a bug report or a custom build request for the Maven Central Email Scraper? Email neurodata.apify@gmail.com.