Maven Central Email Scraper avatar

Maven Central Email Scraper

Pricing

from $2.49 / 1,000 results

Go to Apify Store
Maven Central Email Scraper

Maven Central Email Scraper

Maven Central Email Scraper SD - Maven Central Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from Maven Central results by keyword, location and email domain - Maven Central email extractor.

Pricing

from $2.49 / 1,000 results

Rating

0.0

(0)

Developer

Neuro Scraper

Neuro Scraper

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

7 days ago

Last modified

Categories

Share

Maven Central Email Scraper — project contacts from indexed artifact listings

Maven Central Email Scraper collects publicly indexed contact emails that appear alongside Java and JVM artifact listings in Google's search index. You supply keywords, the email domains you care about and an optional location, and it returns a deduplicated dataset of project contacts.

Read this before you run it: Maven listings identify artifacts and organisations rather than people, so rows usually carry a project contact instead of a personal handle. In measured runs the Maven Central Email Scraper resolved account identity on only 1 row out of 10, and that is the honest expectation to plan around.

That is not a defect in the Actor — it is how the JVM ecosystem works. A published artifact is identified by a groupId and artifactId, and the contact printed next to it is very often a mailing list, an organisation address or a project inbox rather than an individual developer's.

Unlike most Actors in this family, the Maven Central Email Scraper searches three domains rather than one: mvnrepository.com, central.sonatype.com and search.maven.org. Those three sites index the same repository from different angles, and each renders project metadata a little differently.

The Maven Central Email Scraper does not use any Maven or Sonatype API, does not download JARs and does not read pom.xml or any other build file. Everything it returns comes from Google result blocks — titles, snippets and site labels — fetched through the Apify GOOGLE_SERP proxy.

Who the Maven Central Email Scraper is for

Java tooling vendors, security researchers doing coordinated disclosure on JVM libraries, supply-chain analysts, enterprise open-source programme offices and consultancies selling into JVM shops all want the same thing: a way to reach the organisation behind a published artifact.

If you also work in other package ecosystems, the Maven Central Email Scraper pairs directly with the npm Email Scraper and the PyPI Email Scraper, both of which yield personal maintainer contacts far more often.

Key features of the Maven Central Email Scraper

Everything the Actor can do is in this table. If a capability is not listed, the Maven Central Email Scraper does not have it.

FeatureWhat it does
Three-domain site: targetingQueries are scoped to mvnrepository.com, central.sonatype.com and search.maven.org
Artifact URL buildingConstructs profileUrl as https://mvnrepository.com/artifact/{username} when a coordinate-style handle is exposed
Query expansionRuns each keyword as a base query, a quoted phrase, an intitle: query and one variant per query modifier
Domain filteringKeeps only emails ending in the domains you list, with boundary-correct matching
Obfuscation handlingUnderstands name [at] domain [dot] com, name (at) domain, name @ domain.com, domain .com, zero-width characters and the full-width @
Junk filterRejects placeholders such as email@, yourname@, test@, xxx@ and single-character locals
Soft-wrap repairDrops a hit that is only the tail of another email in the same result block
Global deduplicationOne row per unique email address across all three domains, every query and every page
Truncation flagSets possiblyTruncated: true when Google's snippet ellipsis touched the address
Structural parsingLocates the <h3> title then the smallest surrounding block, rather than depending on Google's CSS class names
Whole-page fallbackA markup change degrades the run to "emails without account details" instead of "no emails"
Block detectionCAPTCHA, "unusual traffic" and consent pages are detected and retried, not counted as empty
Retries and backoffUp to 3 attempts per page with exponential backoff and a fresh proxy session per request
Requeue of failuresBlocked or failed queries are re-queued once at the end of the run
Resumable stateProgress lives in the key-value store keyed by a hash of your input, saved on PERSIST_STATE, MIGRATING and ABORTING
Concurrency controlAn asyncio worker pool with a shared stop signal on maxEmails
Streaming outputEvery lead is pushed to the dataset the moment it is found

How the Maven Central Email Scraper works

Six steps, all of them visible in the run log.

  1. The Maven Central Email Scraper reads your input: keywords, location, email domains and limits.
  2. It builds Google queries with the site: operator against each of its three domains, for example site:mvnrepository.com json parser contact "@gmail.com".
  3. It fetches Google result pages asynchronously with aiohttp through the Apify GOOGLE_SERP proxy.
  4. It parses each result block structurally — find the <h3>, then take the smallest block containing it.
  5. It extracts email addresses from that block's text with a domain-filtered regular expression.
  6. It deduplicates globally and pushes each new lead straight into the Apify dataset.

There is no browser, no JavaScript rendering, no authentication and no cookies. The Maven Central Email Scraper only ever sees what Google already shows the public.

Query expansion matters more here than almost anywhere else. Google caps a single query at roughly 300 results, and with contact addresses appearing sparsely across artifact listings, the extra phrasings are often the difference between an empty dataset and a usable one.

Base queries run first, so the strongest matches land before the long tail. Because three domains are in play, expect the Maven Central Email Scraper to spend a noticeable share of a run on the two Sonatype-operated sites.

Input fields

Every input field of the Maven Central Email Scraper is listed below exactly as defined in the Actor's input schema. Only keywords is required.

FieldTypeDefaultMeaning
keywordsarray (required)["java library", "artifact"]Search terms describing the Maven Central listings you want
locationstring""Optional location phrase added to every query
customDomainsarray["@gmail.com", "@yahoo.com"]Only emails on these domains are kept; the @ is optional
maxEmailsinteger 1–1000020Stop after this many unique emails
countryCodestring""Two-letter country for the search proxy (US, GB, DE…)
expandQueriesbooleantrueSearch each keyword × domain pair in several phrasings
queryModifiersarray["email", "contact", "maintainer", "author", "support"]Extra words combined with each keyword when expansion is on
maxPagesPerQueryinteger 1–5030Page cap per query
maxConcurrencyinteger 1–205Parallel queries

The queryModifiers default is tuned for Maven Central. contact, maintainer and author are the words that appear near a published address on an artifact page, so they carry most of the weight in a run.

Example input for the Maven Central Email Scraper

{
"keywords": ["json parser", "spring boot starter", "kotlin coroutines library"],
"location": "",
"customDomains": ["@gmail.com", "@yahoo.com"],
"maxEmails": 300,
"countryCode": "US",
"expandQueries": true,
"queryModifiers": ["email", "contact", "maintainer", "author", "support"],
"maxPagesPerQuery": 30,
"maxConcurrency": 5
}

Describe the artifact, not a person. logging framework, jdbc driver, gradle plugin and protobuf runtime match how artifact pages present themselves far better than any job title.

Output fields

Every dataset item produced by the Maven Central Email Scraper carries all fourteen fields below. Nothing is omitted; fields with no value are empty or null.

FieldMeaning
networkPlatform name
keywordThe keyword that produced the lead
queryThe exact Google query used
titleRaw result title
accountNameAccount label Google prints — on Maven Central this is normally an artifact or organisation
fullNameDisplay name parsed from a profile-style title; empty for other result types
usernameA coordinate-style handle when Google exposes one; otherwise null — frequently null here
profileUrlhttps://mvnrepository.com/artifact/{username} when a handle is known; otherwise empty
urlDirect listing link when exposed, else the profile URL
descriptionArtifact snippet, cleaned of labels and counters
emailLower-cased email address
emailDomainThe matched domain, e.g. @gmail.com
possiblyTruncatedtrue when Google's snippet ellipsis touched the email — verify before sending
foundAtISO 8601 UTC timestamp

Example output from the Maven Central Email Scraper

The first row below is the good case. The second is the common case: a real, usable project address with no handle attached.

[
{
"network": "Maven Central",
"keyword": "json parser",
"query": "site:mvnrepository.com json parser contact \"@gmail.com\"",
"title": "Maven Repository: io.harborline » harborline-json",
"accountName": "io.harborline",
"fullName": "",
"username": "io.harborline",
"profileUrl": "https://mvnrepository.com/artifact/io.harborline",
"url": "https://mvnrepository.com/artifact/io.harborline/harborline-json",
"description": "Streaming JSON parser for the JVM. Project contact: harborline.oss@gmail.com",
"email": "harborline.oss@gmail.com",
"emailDomain": "@gmail.com",
"possiblyTruncated": false,
"foundAt": "2026-08-31T13:11:26Z"
},
{
"network": "Maven Central",
"keyword": "spring boot starter",
"query": "site:central.sonatype.com \"spring boot starter\" maintainer \"@yahoo.com\"",
"title": "Maven Central: northgate-metrics-starter",
"accountName": "northgate-metrics-starter",
"fullName": "",
"username": null,
"profileUrl": "",
"url": "https://central.sonatype.com/artifact/com.northgate/northgate-metrics-starter",
"description": "Metrics auto-configuration starter. Developer mailing list: northgate.dev.list@yahoo.com",
"email": "northgate.dev.list@yahoo.com",
"emailDomain": "@yahoo.com",
"possiblyTruncated": true,
"foundAt": "2026-08-31T13:14:02Z"
}
]

Rows like the second one are the norm rather than the exception on this platform. They are still useful: email, accountName, url and description are all populated by the Maven Central Email Scraper, you simply do not get a handle.

Use cases for the Maven Central Email Scraper

Use caseHow the Maven Central Email Scraper helps
Coordinated vulnerability disclosureReach the project inbox published next to a JVM artifact
Java tooling and platform salesTarget organisations publishing artifacts in the space your product serves
Supply-chain and dependency researchMap which organisations publish which corner of the repository
Enterprise open-source programme outreachContact the teams behind libraries your organisation depends on
Library adoption or handoverReach the owning organisation about an artifact you would like to maintain
Sponsorship and funding programmesIdentify JVM projects with a published contact channel
Licensing and compliance queriesFind the correspondence address for an artifact's owning organisation
CRM enrichmentAttach a published project contact to an organisation you already track

Because most JVM libraries are developed on a public forge, the Maven Central Email Scraper works well next to the GitHub Email Scraper, which is far more likely to give you a named individual.

Tips to get more from the Maven Central Email Scraper

Feed the Maven Central Email Scraper artifact functions, not people. connection pool, bytecode manipulation, test containers, xml binding and metrics exporter all appear in indexed artifact titles; job titles essentially never do.

Widen customDomains aggressively. Because JVM projects usually publish an organisational or mailing-list address, the default @gmail.com and @yahoo.com pair is often the tightest constraint on a Maven Central Email Scraper run — add the project and company domains you actually care about.

Never turn expandQueries off here. Contact addresses are sparse across artifact listings, so the modifier variants are doing most of the work in a Maven Central Email Scraper run.

Expect to iterate. Several narrow Maven Central Email Scraper runs with different artifact vocabularies will out-perform one broad run, and because deduplication is per run you can merge the datasets afterwards.

Do not expect username on most rows. Plan your downstream pipeline around email, accountName, url and description, and treat a populated handle as a bonus.

Limitations you should know before running the Maven Central Email Scraper

Maven listings identify artifacts and organisations rather than people, so rows usually carry a project contact instead of a personal handle. Measured account-identity resolution was 1 out of 10 rows. This is the dominant limitation and it will not change with better keywords.

The Maven Central Email Scraper only finds addresses that are publicly visible in Google's index. An address that exists only inside a build file will never appear.

Google caps a single query at roughly 300 results. Query expansion mitigates that ceiling across three domains; it does not remove it.

username and profileUrl are populated only when Google's result exposes a coordinate-style handle. Most rows arrive with username: null and profileUrl: "", keeping accountName, email and url. That is a Google limitation, not a bug.

possiblyTruncated: true means Google's snippet ellipsis touched the address and it may be incomplete. Verify before sending.

The Actor requires the Apify GOOGLE_SERP proxy and cannot run without Apify proxy credentials. Free Apify plans are capped at 100 emails per run; paid plans are uncapped.

Results vary with keywords, domains, country and timing, and no volume is guaranteed. The Maven Central Email Scraper reads no repository API and no build file, so pom.xml contents are outside its reach.

Responsible use

Contact addresses published next to a Maven artifact exist so that downstream users, security researchers and packagers can correspond with the project. Use them for that kind of correspondence.

Respect GDPR and equivalent local rules, honour opt-outs immediately and follow each project's stated contact norms. Many of these are shared mailing lists, so unrelated marketing is visible to the whole project and reflects badly on you.

The Maven Central Email Scraper is a discovery tool, not a sending tool. What you do with the resulting dataset is entirely your responsibility.

Maven Central Email Scraper FAQ

Does the Maven Central Email Scraper use a repository API or read pom.xml?

No. It uses no Maven or Sonatype API, downloads no JARs and reads no build file. Every field comes from Google search result blocks fetched through the Apify GOOGLE_SERP proxy.

Three: mvnrepository.com, central.sonatype.com and search.maven.org. The Maven Central Email Scraper deduplicates across all three, so a project indexed on more than one site still produces a single row per unique email.

Why is username null on most rows?

Because Maven listings identify artifacts and organisations rather than people, so Google's result often exposes no coordinate-style handle at all. Those rows still carry accountName, email, url and description.

How many rows get a resolved handle?

Roughly one in ten in measured runs. Treat a populated username and profileUrl as a bonus rather than something to build a pipeline on.

Are the emails the Maven Central Email Scraper returns personal addresses?

Usually not. Expect project inboxes, mailing lists and organisational addresses more often than an individual developer's personal email.

How is profileUrl built when a handle exists?

As https://mvnrepository.com/artifact/{username}, using the coordinate-style handle Google exposed.

Which email domains does the Maven Central Email Scraper collect?

Only the domains in customDomains, defaulting to @gmail.com and @yahoo.com. Matching is boundary-correct, so @gmail.com never matches inside @gmail.company or @gmail.com.br.

Does it handle obfuscated addresses?

Yes. The Maven Central Email Scraper normalises name [at] domain [dot] com, name (at) domain, name @ domain.com, domain .com, zero-width characters and the full-width @ before filtering by domain.

What is possiblyTruncated for?

It flags rows where Google's snippet ellipsis touched the email, so the address may be cut off. Verify those before sending.

Do I need a proxy?

Yes. The Maven Central Email Scraper runs on the Apify GOOGLE_SERP proxy and cannot operate without Apify proxy credentials.

Can I resume an interrupted run?

Yes. Progress is stored in the key-value store keyed by a hash of your input, with throttled saves that also fire on PERSIST_STATE, MIGRATING and ABORTING.

What happens if Google changes its HTML?

Parsing is structural rather than class-name based, and a whole-page fallback parser exists. A layout change degrades the Maven Central Email Scraper to "emails without account details" rather than "no emails at all".

The Maven Central Email Scraper is one of 32 Developer & Technology email Actors that share the same engine and differ only in the platform they target.

ActorWhat it collects
Maven Central Email and Phone Number ScraperEmails and phone numbers from Maven Central
Maven Central Phone Number ScraperPublic phone numbers from Maven Central
App Store Email ScraperPublic contact emails from App Store
Atlassian Marketplace Email ScraperPublic contact emails from Atlassian Marketplace
Bitbucket Email ScraperPublic contact emails from Bitbucket
Chrome Web Store Email ScraperPublic contact emails from Chrome Web Store
CodePen Email ScraperPublic contact emails from CodePen
Confluence Email ScraperPublic contact emails from Confluence
Dev.to Email ScraperPublic contact emails from DEV Community
Docker Hub Email ScraperPublic contact emails from Docker Hub
Figma Community Email ScraperPublic contact emails from Figma Community
Firefox Add-ons Email ScraperPublic contact emails from Firefox Add-ons
GitHub Email ScraperPublic contact emails from GitHub
GitLab Email ScraperPublic contact emails from GitLab
Google Play Email ScraperPublic contact emails from Google Play
Hashnode Email ScraperPublic contact emails from Hashnode
HubSpot Marketplace Email ScraperPublic contact emails from HubSpot Marketplace
Hugging Face Email ScraperPublic contact emails from Hugging Face
Jira Email ScraperPublic contact emails from Jira
Microsoft AppSource Email ScraperPublic contact emails from Microsoft AppSource
Salesforce AppExchange Email ScraperPublic contact emails from Salesforce AppExchange
Shopify App Store Email ScraperPublic contact emails from Shopify App Store
Slack App Directory Email ScraperPublic contact emails from Slack App Directory
SourceForge Email ScraperPublic contact emails from SourceForge
Stack Overflow Email ScraperPublic contact emails from Stack Overflow
Unity Asset Store Email ScraperPublic contact emails from Unity Asset Store
Unreal Engine Marketplace Email ScraperPublic contact emails from Unreal Engine Marketplace
WordPress Plugin Directory Email ScraperPublic contact emails from WordPress Plugin Directory
WordPress Theme Directory Email ScraperPublic contact emails from WordPress Theme Directory
Zapier App Directory Email ScraperPublic contact emails from Zapier App Directory
App Store Email and Phone Number ScraperEmails and phone numbers from App Store
Atlassian Marketplace Email and Phone Number ScraperEmails and phone numbers from Atlassian Marketplace
Bitbucket Email and Phone Number ScraperEmails and phone numbers from Bitbucket
Chrome Web Store Email and Phone Number ScraperEmails and phone numbers from Chrome Web Store
CodePen Email and Phone Number ScraperEmails and phone numbers from CodePen
Confluence Email and Phone Number ScraperEmails and phone numbers from Confluence
DEV Community Email and Phone Number ScraperEmails and phone numbers from DEV Community
Docker Hub Email and Phone Number ScraperEmails and phone numbers from Docker Hub
Figma Community Email and Phone Number ScraperEmails and phone numbers from Figma Community
Firefox Add-ons Email and Phone Number ScraperEmails and phone numbers from Firefox Add-ons
GitHub Email and Phone Number ScraperEmails and phone numbers from GitHub
GitLab Email and Phone Number ScraperEmails and phone numbers from GitLab

Leave a review

If the Maven Central Email Scraper saved you time, please leave a star rating and a short review on the Actor page.

Reviews are how other buyers judge whether a tool works, and they tell us which features to build next.

If something did not work, email neurodata.apify@gmail.com instead - bugs get fixed faster than they get complained about.

Support

Questions, a bug report or a custom build request for the Maven Central Email Scraper? Email neurodata.apify@gmail.com.