SourceForge Email Scraper avatar

SourceForge Email Scraper

Pricing

from $2.49 / 1,000 results

Go to Apify Store
SourceForge Email Scraper

SourceForge Email Scraper

SourceForge Email Scraper SD - SourceForge Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from SourceForge results by keyword, location and email domain - SourceForge email extractor.

Pricing

from $2.49 / 1,000 results

Rating

0.0

(0)

Developer

Leads Scraper

Leads Scraper

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

16 days ago

Last modified

Categories

Share

SourceForge Email Scraper - maintainer contacts from project and user pages

The SourceForge Email Scraper is an Apify Actor that collects publicly indexed contact emails from SourceForge project pages, user profiles, wikis and ticket threads. Give it keywords, an optional location and your email domains, and it returns a structured lead dataset.

SourceForge is the long memory of open source. Projects that predate GitHub are still hosted, still downloaded and still maintained there, and many of them publish a maintainer address in plain text on the project page.

That is exactly what the SourceForge Email Scraper is for: reaching the people behind mature, long-lived open-source software that never migrated anywhere else.

What the SourceForge Email Scraper reads

It reads only Google's public index of sourceforge.net. It builds site: queries, fetches result pages through the Apify GOOGLE_SERP proxy, and extracts emails from the result titles and snippets.

It does not read git or Subversion history, clone repositories, or use any SourceForge API. There is no login, no browser, no JavaScript rendering and no cookies in the pipeline.

Addresses surface from project summary pages, README and documentation files, wiki pages, ticket and discussion threads, and user profile pages under /u/ - wherever a maintainer chose to publish one and Google indexed it.

Who the SourceForge Email Scraper is built for

Security researchers contacting maintainers of dependencies, developer relations teams reaching legacy-project owners, agencies doing developer lead generation in niche verticals, and anyone who needs to talk to the person still shipping a twenty-year-old utility.

If you have hunted through a SourceForge project page for a way to email its owner, this Actor does that at scale and returns a spreadsheet.


Key features of the SourceForge Email Scraper

FeatureWhat it does
Google site: dorkingQueries sourceforge.net through the Apify GOOGLE_SERP proxy
Query expansionBase, quoted and intitle: variants plus one variant per query modifier; base queries run first
Domain-filtered extractionKeeps only emails ending in your customDomains values
Global deduplicationOne row per unique email address across every query and page
Obfuscation handlingUnderstands name [at] domain [dot] com, name (at) domain, name @ domain.com, domain .com, zero-width characters and the full-width @ - common in older project pages written to dodge spam bots
Junk filterRejects placeholders such as email@, yourname@, test@, xxx@ and single-character locals
Boundary-correct matching@gmail.com never matches inside @gmail.company or @gmail.com.br
Soft-wrap repairDiscards a hit that is only the tail of another email in the same result block
Structural parsingLocates the <h3> title then the smallest surrounding block, not Google's CSS class names
Whole-page fallbackA Google markup change degrades the run to "emails without account details", not "no emails"
Block detectionCAPTCHA, "unusual traffic" and consent pages are detected and retried rather than treated as empty
Retries and backoffUp to 3 attempts per page, exponential backoff, a fresh proxy session per request
Requeue of failuresBlocked or failed queries are re-queued once at the end of the run
Async concurrencyAn asyncio worker pool with a shared stop signal on maxEmails
Resumable stateKey-value-store progress keyed by an input hash, with throttled saves plus saves on PERSIST_STATE, MIGRATING and ABORTING
Streaming outputLeads reach the dataset as they are found, so an aborted run keeps what it collected
Run summaryLogs pages fetched, blocked pages, retries and emails per page

The obfuscation handling is unusually relevant on SourceForge. Older project pages routinely write maintainer [at] example [dot] com, and the SourceForge Email Scraper normalises those into real addresses.


How the SourceForge Email Scraper works

The SourceForge Email Scraper pipeline is six steps, with no authentication anywhere in it.

  1. It reads your input: keywords, location, email domains and limits.
  2. It builds Google queries with the site: operator, for example site:sourceforge.net maintainer contact "@gmail.com".
  3. It fetches Google result pages asynchronously with aiohttp through the Apify GOOGLE_SERP proxy.
  4. It parses each result block structurally, finding the <h3> title and the smallest block around it.
  5. It extracts email addresses from the block text using a domain-filtered regex.
  6. It deduplicates globally and pushes each lead straight into the Apify dataset.

Why query expansion matters more here

SourceForge is a single domain, so there is no second site to spread queries across the way github.io complements github.com.

Google caps one query at roughly 300 results. With only one domain to target, query expansion is the main lever the SourceForge Email Scraper has for increasing total coverage.

The default queryModifiers - email, contact, maintainer, author, support - match how SourceForge project pages label their contact sections.

What the SourceForge Email Scraper does not do

No SourceForge login, no API, no repository or SVN checkout, no commit-history parsing, no JavaScript rendering. It is an independent Apify Actor, not affiliated with or endorsed by SourceForge or Slashdot Media.


SourceForge Email Scraper input fields

Only keywords is required by the SourceForge Email Scraper. All defaults below come from the Actor's shipped input schema.

FieldTypeDefaultMeaning
keywordsarray (required)["maintainer", "open source"]Search terms describing the SourceForge accounts and projects you want (niche, job title, technology)
locationstring""Optional location phrase added to every query
customDomainsarray["@gmail.com", "@yahoo.com"]Only emails on these domains are kept; the leading @ is optional
maxEmailsinteger 1-1000020Stop after this many unique emails
countryCodestring""Two-letter country for the search proxy (US, GB, DE...)
expandQueriesbooleantrueSearch each keyword x domain pair in several phrasings
queryModifiersarray["email", "contact", "maintainer", "author", "support"]Extra words combined with each keyword when expansion is on
maxPagesPerQueryinteger 1-5030Page cap per query
maxConcurrencyinteger 1-205Parallel queries

Example SourceForge Email Scraper input

{
"keywords": ["cnc controller", "scientific computing", "embedded toolchain"],
"location": "",
"customDomains": ["@gmail.com", "@yahoo.com"],
"maxEmails": 300,
"countryCode": "US",
"expandQueries": true,
"queryModifiers": ["email", "contact", "maintainer", "author", "support"],
"maxPagesPerQuery": 30,
"maxConcurrency": 5
}

Tuning the SourceForge Email Scraper

Domain-specific project vocabulary works best in the SourceForge Email Scraper. plc simulator, astronomy imaging, serial terminal and similar phrases match how SourceForge projects actually describe themselves.

Because there is only one site to query, add email domains generously - @yahoo.com, @hotmail.com and older free providers are well represented in this population.

Keep expandQueries on. With a single domain to target, disabling it is the fastest way to cap the SourceForge Email Scraper at a few hundred results.


SourceForge Email Scraper output fields

Every dataset item the SourceForge Email Scraper writes has all 14 fields. Values are empty or null when SourceForge did not expose them in Google's result - they are never dropped.

FieldMeaning
networkPlatform name
keywordThe keyword that produced the lead
queryThe exact Google query used
titleRaw result title
accountNameAccount label Google prints - a user handle on profile pages, a project name on project pages
fullNameDisplay name parsed from a profile-style title; empty for project and doc pages
usernameURL-safe SourceForge handle when one is exposed; otherwise null
profileUrlCanonical https://sourceforge.net/u/{username}/profile/ URL when a handle is known; otherwise empty
urlDirect platform link when exposed, else the profile URL
descriptionSnippet text, cleaned of labels and engagement counters
emailLower-cased email address
emailDomainThe matched domain, for example @gmail.com
possiblyTruncatedtrue when Google's snippet ellipsis touched the email - verify before sending
foundAtISO 8601 UTC timestamp

Example SourceForge Email Scraper output

[
{
"network": "SourceForge",
"keyword": "cnc controller",
"query": "site:sourceforge.net cnc controller maintainer \"@gmail.com\"",
"title": "mgrieder - SourceForge",
"accountName": "mgrieder",
"fullName": "Martin Grieder",
"username": "mgrieder",
"profileUrl": "https://sourceforge.net/u/mgrieder/profile/",
"url": "https://sourceforge.net/u/mgrieder/profile/",
"description": "Maintainer of a stepper motion control suite. Patches and questions to martin.grieder.cnc@gmail.com",
"email": "martin.grieder.cnc@gmail.com",
"emailDomain": "@gmail.com",
"possiblyTruncated": false,
"foundAt": "2026-08-31T13:10:33.402Z"
},
{
"network": "SourceForge",
"keyword": "scientific computing",
"query": "site:sourceforge.net scientific computing contact \"@yahoo.com\"",
"title": "AstroReduce - Browse Files at SourceForge.net",
"accountName": "AstroReduce",
"fullName": "",
"username": null,
"profileUrl": "",
"url": "https://sourceforge.net/projects/astroreduce/files/",
"description": "Bug reports to astroreduce.project@yahoo.com. Release notes for 4.2 are in the wiki ...",
"email": "astroreduce.project@yahoo.com",
"emailDomain": "@yahoo.com",
"possiblyTruncated": true,
"foundAt": "2026-08-31T13:10:51.860Z"
},
{
"network": "SourceForge",
"keyword": "embedded toolchain",
"query": "site:sourceforge.net intitle:\"embedded toolchain\" support \"@gmail.com\"",
"title": "Ticket #148: cross compiler build fails on ARM",
"accountName": "",
"fullName": "",
"username": null,
"profileUrl": "",
"url": "https://sourceforge.net/p/xtool-embedded/tickets/148/",
"description": "Reported by the toolchain author, reachable at xtool.embedded.dev@gmail.com for build issues",
"email": "xtool.embedded.dev@gmail.com",
"emailDomain": "@gmail.com",
"possiblyTruncated": false,
"foundAt": "2026-08-31T13:11:24.517Z"
}
]

This is a representative mix of SourceForge Email Scraper output. The first row is a user profile with a full identity, the second a project page with a project name but no handle, the third a ticket thread with an email and no account details at all.


Use cases for the SourceForge Email Scraper

Use caseHow the SourceForge Email Scraper helps
Security disclosureReach maintainers of legacy dependencies who have no SECURITY.md and no issue tracker you can use
Open-source outreachContact project owners about forks, ports, packaging or maintenance handover
Developer relationsFind owners of long-lived projects in your category before a launch or migration campaign
Developer lead generationBuild a niche list of technical project owners in vertical software
Vendor and partner discoveryIdentify who still maintains tooling your customers depend on
Migration and modernisation offersApproach owners of projects that would benefit from a newer platform
Academic and research outreachReach authors of scientific and engineering utilities published on SourceForge
Ecosystem and archival researchMeasure how many projects in a niche still publish a working contact address

The security disclosure angle

Plenty of software in production dependency trees still lives on SourceForge, and much of it has no modern reporting channel. The SourceForge Email Scraper often finds the only working contact route.

Search your dependency names as keywords, set the domains you expect, and check possiblyTruncated before you send anything sensitive.

Combining with sibling Actors

Run the SourceForge Email Scraper together with the GitHub Email Scraper and the PyPI Email Scraper to catch maintainers who publish source in one place and packages in another.


Expected results from the SourceForge Email Scraper

In live test runs against Google, about 4 out of 10 parsed results carried an account identity - a handle, a display name, or both. Output is a genuine mix of user pages and project pages.

User profile pages under /u/ resolve cleanly to a handle and a profile URL. Project pages, download listings, wikis and tickets typically give you a project name instead, and sometimes nothing but the email itself.

Every SourceForge Email Scraper row still carries email, emailDomain, url, title, description and query, which is enough to qualify a lead. Volume depends on keywords and domains; no yield is guaranteed.


Limitations of the SourceForge Email Scraper

  • Publicly indexed emails only. An address is findable only when it is already visible in Google's index. Anything behind a login or never published is unreachable.
  • No repository history, no API. The SourceForge Email Scraper does not read git or SVN history, clone repositories, or call any SourceForge API. Only Google result titles and snippets are parsed.
  • Google's ~300-result cap. One query returns roughly 300 results at most. Because SourceForge is a single domain, expandQueries is the main way to go beyond that.
  • possiblyTruncated. When true, Google's snippet ellipsis touched the email and it may be cut off. Verify those rows before sending.
  • username and profileUrl are often empty. They are populated only when SourceForge exposes a handle in the result. Project, download, wiki and ticket pages usually show a project name instead, leaving accountName filled but username null.
  • Stale contacts. Many SourceForge projects are old. An address published a decade ago may no longer be monitored - treat bounces as expected.
  • Apify GOOGLE_SERP proxy required. The Actor cannot run without Apify proxy credentials.
  • Free plan cap. Free Apify plans are limited to 100 emails per run; paid plans are uncapped.
  • Variable results. Output depends on keywords, domains and location, and Google's index changes over time.

Responsible use

Maintainer emails the SourceForge Email Scraper returns were published for project correspondence - bug reports, patches, packaging questions, security disclosures. That is the spirit in which to use them.

If you contact people in the EU or UK, GDPR applies: have a lawful basis, identify yourself, say where you found the address, and honour opt-outs immediately.

Many of these projects are maintained by one person in their spare time. Be specific, be brief, and do not send a maintainer a marketing sequence.


SourceForge Email Scraper FAQ

Does the SourceForge Email Scraper read repository history or use an API?

No. It does not read git or SVN history, clone repositories, or call any SourceForge API. It parses Google search results for already-indexed public pages on sourceforge.net.

Do I need a SourceForge account?

No. There is no authentication, no browser, no JavaScript rendering and no cookies. The only credential involved is your Apify account's GOOGLE_SERP proxy access.

Where does the SourceForge Email Scraper get the emails?

From Google's result titles and snippets: project summary pages, README and documentation files, wiki pages, ticket and discussion threads, and user profile pages under /u/.

How often do SourceForge Email Scraper rows include a profile URL?

In live testing, about 4 in 10 parsed results carried an account identity. User profile pages resolve well; project, download and ticket pages usually do not.

Why is username sometimes null?

Because Google exposed a project name rather than a SourceForge handle for that result. The row is still a usable lead with email, accountName, url and description.

What does possiblyTruncated: true mean?

Google's snippet ellipsis touched the address, so it may be incomplete. Verify those rows before you use them.

Can the SourceForge Email Scraper filter to specific domains?

Yes. Set customDomains to what you want, for example ["@gmail.com", "@yahoo.com", "@acme.io"]. The leading @ is optional and only matching addresses are kept.

Does the SourceForge Email Scraper decode maintainer [at] example [dot] com?

Yes. That style is common on older SourceForge pages, and the extractor normalises [at], (at), spaced @, spaced .com, zero-width characters and the full-width @.

How many emails can one SourceForge Email Scraper run return?

maxEmails accepts 1 to 10000 and defaults to 20. Free Apify plans are capped at 100 emails per run; paid plans are uncapped.

How do I target one country?

Set countryCode to a two-letter code such as US, GB or DE, and put a city or region in location so it is appended to every query.

What happens if a SourceForge Email Scraper run is interrupted or migrated?

Progress is stored in the key-value store keyed by a hash of your input, with saves on PERSIST_STATE, MIGRATING and ABORTING. Leads already pushed to the dataset are kept.

Is the SourceForge Email Scraper affiliated with SourceForge?

No. It is an independent Apify Actor, not supported, endorsed or affiliated with SourceForge or Slashdot Media.


The SourceForge Email Scraper belongs to a Developer & Technology family of Apify Actors that apply the same method to different platforms. Run several to cover an ecosystem instead of one site.

ActorWhat it collects
SourceForge Email and Phone Number ScraperEmails and phone numbers from SourceForge
SourceForge Phone Number ScraperPublic phone numbers from SourceForge
App Store Email ScraperPublic contact emails from App Store
Atlassian Marketplace Email ScraperPublic contact emails from Atlassian Marketplace
Bitbucket Email ScraperPublic contact emails from Bitbucket
Chrome Web Store Email ScraperPublic contact emails from Chrome Web Store
CodePen Email ScraperPublic contact emails from CodePen
Confluence Email ScraperPublic contact emails from Confluence
Dev.to Email ScraperPublic contact emails from DEV Community
Docker Hub Email ScraperPublic contact emails from Docker Hub
Figma Community Email ScraperPublic contact emails from Figma Community
Firefox Add-ons Email ScraperPublic contact emails from Firefox Add-ons
GitHub Email ScraperPublic contact emails from GitHub
GitLab Email ScraperPublic contact emails from GitLab
Google Play Email ScraperPublic contact emails from Google Play
Hashnode Email ScraperPublic contact emails from Hashnode
HubSpot Marketplace Email ScraperPublic contact emails from HubSpot Marketplace
Hugging Face Email ScraperPublic contact emails from Hugging Face
Jira Email ScraperPublic contact emails from Jira
Maven Central Email ScraperPublic contact emails from Maven Central
Microsoft AppSource Email ScraperPublic contact emails from Microsoft AppSource
Salesforce AppExchange Email ScraperPublic contact emails from Salesforce AppExchange
Shopify App Store Email ScraperPublic contact emails from Shopify App Store
Slack App Directory Email ScraperPublic contact emails from Slack App Directory
Stack Overflow Email ScraperPublic contact emails from Stack Overflow
Unity Asset Store Email ScraperPublic contact emails from Unity Asset Store
Unreal Engine Marketplace Email ScraperPublic contact emails from Unreal Engine Marketplace
WordPress Plugin Directory Email ScraperPublic contact emails from WordPress Plugin Directory
WordPress Theme Directory Email ScraperPublic contact emails from WordPress Theme Directory
Zapier App Directory Email ScraperPublic contact emails from Zapier App Directory
App Store Email and Phone Number ScraperEmails and phone numbers from App Store
Atlassian Marketplace Email and Phone Number ScraperEmails and phone numbers from Atlassian Marketplace
Bitbucket Email and Phone Number ScraperEmails and phone numbers from Bitbucket
Chrome Web Store Email and Phone Number ScraperEmails and phone numbers from Chrome Web Store
CodePen Email and Phone Number ScraperEmails and phone numbers from CodePen
Confluence Email and Phone Number ScraperEmails and phone numbers from Confluence
DEV Community Email and Phone Number ScraperEmails and phone numbers from DEV Community
Docker Hub Email and Phone Number ScraperEmails and phone numbers from Docker Hub
Figma Community Email and Phone Number ScraperEmails and phone numbers from Figma Community
Firefox Add-ons Email and Phone Number ScraperEmails and phone numbers from Firefox Add-ons
GitHub Email and Phone Number ScraperEmails and phone numbers from GitHub
GitLab Email and Phone Number ScraperEmails and phone numbers from GitLab

Leave a review

If the SourceForge Email Scraper saved you time, please leave a star rating and a short review on the Actor page.

Reviews are how other buyers judge whether a tool works, and they tell us which features to build next.

If something did not work, email neurodata.apify@gmail.com instead - bugs get fixed faster than they get complained about.

Support

Questions, bug reports, or a custom build of the SourceForge Email Scraper tuned to your own domain list? Email neurodata.apify@gmail.com.