SourceForge Email Scraper
Pricing
from $2.49 / 1,000 results
SourceForge Email Scraper
SourceForge Email Scraper SD - SourceForge Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from SourceForge results by keyword, location and email domain - SourceForge email extractor.
Pricing
from $2.49 / 1,000 results
Rating
0.0
(0)
Developer
Leads Scraper
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
16 days ago
Last modified
Categories
Share
SourceForge Email Scraper - maintainer contacts from project and user pages
The SourceForge Email Scraper is an Apify Actor that collects publicly indexed contact emails from SourceForge project pages, user profiles, wikis and ticket threads. Give it keywords, an optional location and your email domains, and it returns a structured lead dataset.
SourceForge is the long memory of open source. Projects that predate GitHub are still hosted, still downloaded and still maintained there, and many of them publish a maintainer address in plain text on the project page.
That is exactly what the SourceForge Email Scraper is for: reaching the people behind mature, long-lived open-source software that never migrated anywhere else.
What the SourceForge Email Scraper reads
It reads only Google's public index of sourceforge.net. It builds site: queries, fetches result pages through the Apify GOOGLE_SERP proxy, and extracts emails from the result titles and snippets.
It does not read git or Subversion history, clone repositories, or use any SourceForge API. There is no login, no browser, no JavaScript rendering and no cookies in the pipeline.
Addresses surface from project summary pages, README and documentation files, wiki pages, ticket and discussion threads, and user profile pages under /u/ - wherever a maintainer chose to publish one and Google indexed it.
Who the SourceForge Email Scraper is built for
Security researchers contacting maintainers of dependencies, developer relations teams reaching legacy-project owners, agencies doing developer lead generation in niche verticals, and anyone who needs to talk to the person still shipping a twenty-year-old utility.
If you have hunted through a SourceForge project page for a way to email its owner, this Actor does that at scale and returns a spreadsheet.
Key features of the SourceForge Email Scraper
| Feature | What it does |
|---|---|
Google site: dorking | Queries sourceforge.net through the Apify GOOGLE_SERP proxy |
| Query expansion | Base, quoted and intitle: variants plus one variant per query modifier; base queries run first |
| Domain-filtered extraction | Keeps only emails ending in your customDomains values |
| Global deduplication | One row per unique email address across every query and page |
| Obfuscation handling | Understands name [at] domain [dot] com, name (at) domain, name @ domain.com, domain .com, zero-width characters and the full-width @ - common in older project pages written to dodge spam bots |
| Junk filter | Rejects placeholders such as email@, yourname@, test@, xxx@ and single-character locals |
| Boundary-correct matching | @gmail.com never matches inside @gmail.company or @gmail.com.br |
| Soft-wrap repair | Discards a hit that is only the tail of another email in the same result block |
| Structural parsing | Locates the <h3> title then the smallest surrounding block, not Google's CSS class names |
| Whole-page fallback | A Google markup change degrades the run to "emails without account details", not "no emails" |
| Block detection | CAPTCHA, "unusual traffic" and consent pages are detected and retried rather than treated as empty |
| Retries and backoff | Up to 3 attempts per page, exponential backoff, a fresh proxy session per request |
| Requeue of failures | Blocked or failed queries are re-queued once at the end of the run |
| Async concurrency | An asyncio worker pool with a shared stop signal on maxEmails |
| Resumable state | Key-value-store progress keyed by an input hash, with throttled saves plus saves on PERSIST_STATE, MIGRATING and ABORTING |
| Streaming output | Leads reach the dataset as they are found, so an aborted run keeps what it collected |
| Run summary | Logs pages fetched, blocked pages, retries and emails per page |
The obfuscation handling is unusually relevant on SourceForge. Older project pages routinely write maintainer [at] example [dot] com, and the SourceForge Email Scraper normalises those into real addresses.
How the SourceForge Email Scraper works
The SourceForge Email Scraper pipeline is six steps, with no authentication anywhere in it.
- It reads your input: keywords, location, email domains and limits.
- It builds Google queries with the
site:operator, for examplesite:sourceforge.net maintainer contact "@gmail.com". - It fetches Google result pages asynchronously with
aiohttpthrough the Apify GOOGLE_SERP proxy. - It parses each result block structurally, finding the
<h3>title and the smallest block around it. - It extracts email addresses from the block text using a domain-filtered regex.
- It deduplicates globally and pushes each lead straight into the Apify dataset.
Why query expansion matters more here
SourceForge is a single domain, so there is no second site to spread queries across the way github.io complements github.com.
Google caps one query at roughly 300 results. With only one domain to target, query expansion is the main lever the SourceForge Email Scraper has for increasing total coverage.
The default queryModifiers - email, contact, maintainer, author, support - match how SourceForge project pages label their contact sections.
What the SourceForge Email Scraper does not do
No SourceForge login, no API, no repository or SVN checkout, no commit-history parsing, no JavaScript rendering. It is an independent Apify Actor, not affiliated with or endorsed by SourceForge or Slashdot Media.
SourceForge Email Scraper input fields
Only keywords is required by the SourceForge Email Scraper. All defaults below come from the Actor's shipped input schema.
| Field | Type | Default | Meaning |
|---|---|---|---|
keywords | array (required) | ["maintainer", "open source"] | Search terms describing the SourceForge accounts and projects you want (niche, job title, technology) |
location | string | "" | Optional location phrase added to every query |
customDomains | array | ["@gmail.com", "@yahoo.com"] | Only emails on these domains are kept; the leading @ is optional |
maxEmails | integer 1-10000 | 20 | Stop after this many unique emails |
countryCode | string | "" | Two-letter country for the search proxy (US, GB, DE...) |
expandQueries | boolean | true | Search each keyword x domain pair in several phrasings |
queryModifiers | array | ["email", "contact", "maintainer", "author", "support"] | Extra words combined with each keyword when expansion is on |
maxPagesPerQuery | integer 1-50 | 30 | Page cap per query |
maxConcurrency | integer 1-20 | 5 | Parallel queries |
Example SourceForge Email Scraper input
{"keywords": ["cnc controller", "scientific computing", "embedded toolchain"],"location": "","customDomains": ["@gmail.com", "@yahoo.com"],"maxEmails": 300,"countryCode": "US","expandQueries": true,"queryModifiers": ["email", "contact", "maintainer", "author", "support"],"maxPagesPerQuery": 30,"maxConcurrency": 5}
Tuning the SourceForge Email Scraper
Domain-specific project vocabulary works best in the SourceForge Email Scraper. plc simulator, astronomy imaging, serial terminal and similar phrases match how SourceForge projects actually describe themselves.
Because there is only one site to query, add email domains generously - @yahoo.com, @hotmail.com and older free providers are well represented in this population.
Keep expandQueries on. With a single domain to target, disabling it is the fastest way to cap the SourceForge Email Scraper at a few hundred results.
SourceForge Email Scraper output fields
Every dataset item the SourceForge Email Scraper writes has all 14 fields. Values are empty or null when SourceForge did not expose them in Google's result - they are never dropped.
| Field | Meaning |
|---|---|
network | Platform name |
keyword | The keyword that produced the lead |
query | The exact Google query used |
title | Raw result title |
accountName | Account label Google prints - a user handle on profile pages, a project name on project pages |
fullName | Display name parsed from a profile-style title; empty for project and doc pages |
username | URL-safe SourceForge handle when one is exposed; otherwise null |
profileUrl | Canonical https://sourceforge.net/u/{username}/profile/ URL when a handle is known; otherwise empty |
url | Direct platform link when exposed, else the profile URL |
description | Snippet text, cleaned of labels and engagement counters |
email | Lower-cased email address |
emailDomain | The matched domain, for example @gmail.com |
possiblyTruncated | true when Google's snippet ellipsis touched the email - verify before sending |
foundAt | ISO 8601 UTC timestamp |
Example SourceForge Email Scraper output
[{"network": "SourceForge","keyword": "cnc controller","query": "site:sourceforge.net cnc controller maintainer \"@gmail.com\"","title": "mgrieder - SourceForge","accountName": "mgrieder","fullName": "Martin Grieder","username": "mgrieder","profileUrl": "https://sourceforge.net/u/mgrieder/profile/","url": "https://sourceforge.net/u/mgrieder/profile/","description": "Maintainer of a stepper motion control suite. Patches and questions to martin.grieder.cnc@gmail.com","email": "martin.grieder.cnc@gmail.com","emailDomain": "@gmail.com","possiblyTruncated": false,"foundAt": "2026-08-31T13:10:33.402Z"},{"network": "SourceForge","keyword": "scientific computing","query": "site:sourceforge.net scientific computing contact \"@yahoo.com\"","title": "AstroReduce - Browse Files at SourceForge.net","accountName": "AstroReduce","fullName": "","username": null,"profileUrl": "","url": "https://sourceforge.net/projects/astroreduce/files/","description": "Bug reports to astroreduce.project@yahoo.com. Release notes for 4.2 are in the wiki ...","email": "astroreduce.project@yahoo.com","emailDomain": "@yahoo.com","possiblyTruncated": true,"foundAt": "2026-08-31T13:10:51.860Z"},{"network": "SourceForge","keyword": "embedded toolchain","query": "site:sourceforge.net intitle:\"embedded toolchain\" support \"@gmail.com\"","title": "Ticket #148: cross compiler build fails on ARM","accountName": "","fullName": "","username": null,"profileUrl": "","url": "https://sourceforge.net/p/xtool-embedded/tickets/148/","description": "Reported by the toolchain author, reachable at xtool.embedded.dev@gmail.com for build issues","email": "xtool.embedded.dev@gmail.com","emailDomain": "@gmail.com","possiblyTruncated": false,"foundAt": "2026-08-31T13:11:24.517Z"}]
This is a representative mix of SourceForge Email Scraper output. The first row is a user profile with a full identity, the second a project page with a project name but no handle, the third a ticket thread with an email and no account details at all.
Use cases for the SourceForge Email Scraper
| Use case | How the SourceForge Email Scraper helps |
|---|---|
| Security disclosure | Reach maintainers of legacy dependencies who have no SECURITY.md and no issue tracker you can use |
| Open-source outreach | Contact project owners about forks, ports, packaging or maintenance handover |
| Developer relations | Find owners of long-lived projects in your category before a launch or migration campaign |
| Developer lead generation | Build a niche list of technical project owners in vertical software |
| Vendor and partner discovery | Identify who still maintains tooling your customers depend on |
| Migration and modernisation offers | Approach owners of projects that would benefit from a newer platform |
| Academic and research outreach | Reach authors of scientific and engineering utilities published on SourceForge |
| Ecosystem and archival research | Measure how many projects in a niche still publish a working contact address |
The security disclosure angle
Plenty of software in production dependency trees still lives on SourceForge, and much of it has no modern reporting channel. The SourceForge Email Scraper often finds the only working contact route.
Search your dependency names as keywords, set the domains you expect, and check possiblyTruncated before you send anything sensitive.
Combining with sibling Actors
Run the SourceForge Email Scraper together with the GitHub Email Scraper and the PyPI Email Scraper to catch maintainers who publish source in one place and packages in another.
Expected results from the SourceForge Email Scraper
In live test runs against Google, about 4 out of 10 parsed results carried an account identity - a handle, a display name, or both. Output is a genuine mix of user pages and project pages.
User profile pages under /u/ resolve cleanly to a handle and a profile URL. Project pages, download listings, wikis and tickets typically give you a project name instead, and sometimes nothing but the email itself.
Every SourceForge Email Scraper row still carries email, emailDomain, url, title, description and query, which is enough to qualify a lead. Volume depends on keywords and domains; no yield is guaranteed.
Limitations of the SourceForge Email Scraper
- Publicly indexed emails only. An address is findable only when it is already visible in Google's index. Anything behind a login or never published is unreachable.
- No repository history, no API. The SourceForge Email Scraper does not read git or SVN history, clone repositories, or call any SourceForge API. Only Google result titles and snippets are parsed.
- Google's ~300-result cap. One query returns roughly 300 results at most. Because SourceForge is a single domain,
expandQueriesis the main way to go beyond that. possiblyTruncated. Whentrue, Google's snippet ellipsis touched the email and it may be cut off. Verify those rows before sending.usernameandprofileUrlare often empty. They are populated only when SourceForge exposes a handle in the result. Project, download, wiki and ticket pages usually show a project name instead, leavingaccountNamefilled butusernamenull.- Stale contacts. Many SourceForge projects are old. An address published a decade ago may no longer be monitored - treat bounces as expected.
- Apify GOOGLE_SERP proxy required. The Actor cannot run without Apify proxy credentials.
- Free plan cap. Free Apify plans are limited to 100 emails per run; paid plans are uncapped.
- Variable results. Output depends on keywords, domains and location, and Google's index changes over time.
Responsible use
Maintainer emails the SourceForge Email Scraper returns were published for project correspondence - bug reports, patches, packaging questions, security disclosures. That is the spirit in which to use them.
If you contact people in the EU or UK, GDPR applies: have a lawful basis, identify yourself, say where you found the address, and honour opt-outs immediately.
Many of these projects are maintained by one person in their spare time. Be specific, be brief, and do not send a maintainer a marketing sequence.
SourceForge Email Scraper FAQ
Does the SourceForge Email Scraper read repository history or use an API?
No. It does not read git or SVN history, clone repositories, or call any SourceForge API. It parses Google search results for already-indexed public pages on sourceforge.net.
Do I need a SourceForge account?
No. There is no authentication, no browser, no JavaScript rendering and no cookies. The only credential involved is your Apify account's GOOGLE_SERP proxy access.
Where does the SourceForge Email Scraper get the emails?
From Google's result titles and snippets: project summary pages, README and documentation files, wiki pages, ticket and discussion threads, and user profile pages under /u/.
How often do SourceForge Email Scraper rows include a profile URL?
In live testing, about 4 in 10 parsed results carried an account identity. User profile pages resolve well; project, download and ticket pages usually do not.
Why is username sometimes null?
Because Google exposed a project name rather than a SourceForge handle for that result. The row is still a usable lead with email, accountName, url and description.
What does possiblyTruncated: true mean?
Google's snippet ellipsis touched the address, so it may be incomplete. Verify those rows before you use them.
Can the SourceForge Email Scraper filter to specific domains?
Yes. Set customDomains to what you want, for example ["@gmail.com", "@yahoo.com", "@acme.io"]. The leading @ is optional and only matching addresses are kept.
Does the SourceForge Email Scraper decode maintainer [at] example [dot] com?
Yes. That style is common on older SourceForge pages, and the extractor normalises [at], (at), spaced @, spaced .com, zero-width characters and the full-width @.
How many emails can one SourceForge Email Scraper run return?
maxEmails accepts 1 to 10000 and defaults to 20. Free Apify plans are capped at 100 emails per run; paid plans are uncapped.
How do I target one country?
Set countryCode to a two-letter code such as US, GB or DE, and put a city or region in location so it is appended to every query.
What happens if a SourceForge Email Scraper run is interrupted or migrated?
Progress is stored in the key-value store keyed by a hash of your input, with saves on PERSIST_STATE, MIGRATING and ABORTING. Leads already pushed to the dataset are kept.
Is the SourceForge Email Scraper affiliated with SourceForge?
No. It is an independent Apify Actor, not supported, endorsed or affiliated with SourceForge or Slashdot Media.
Related Actors
The SourceForge Email Scraper belongs to a Developer & Technology family of Apify Actors that apply the same method to different platforms. Run several to cover an ecosystem instead of one site.
| Actor | What it collects |
|---|---|
| SourceForge Email and Phone Number Scraper | Emails and phone numbers from SourceForge |
| SourceForge Phone Number Scraper | Public phone numbers from SourceForge |
| App Store Email Scraper | Public contact emails from App Store |
| Atlassian Marketplace Email Scraper | Public contact emails from Atlassian Marketplace |
| Bitbucket Email Scraper | Public contact emails from Bitbucket |
| Chrome Web Store Email Scraper | Public contact emails from Chrome Web Store |
| CodePen Email Scraper | Public contact emails from CodePen |
| Confluence Email Scraper | Public contact emails from Confluence |
| Dev.to Email Scraper | Public contact emails from DEV Community |
| Docker Hub Email Scraper | Public contact emails from Docker Hub |
| Figma Community Email Scraper | Public contact emails from Figma Community |
| Firefox Add-ons Email Scraper | Public contact emails from Firefox Add-ons |
| GitHub Email Scraper | Public contact emails from GitHub |
| GitLab Email Scraper | Public contact emails from GitLab |
| Google Play Email Scraper | Public contact emails from Google Play |
| Hashnode Email Scraper | Public contact emails from Hashnode |
| HubSpot Marketplace Email Scraper | Public contact emails from HubSpot Marketplace |
| Hugging Face Email Scraper | Public contact emails from Hugging Face |
| Jira Email Scraper | Public contact emails from Jira |
| Maven Central Email Scraper | Public contact emails from Maven Central |
| Microsoft AppSource Email Scraper | Public contact emails from Microsoft AppSource |
| Salesforce AppExchange Email Scraper | Public contact emails from Salesforce AppExchange |
| Shopify App Store Email Scraper | Public contact emails from Shopify App Store |
| Slack App Directory Email Scraper | Public contact emails from Slack App Directory |
| Stack Overflow Email Scraper | Public contact emails from Stack Overflow |
| Unity Asset Store Email Scraper | Public contact emails from Unity Asset Store |
| Unreal Engine Marketplace Email Scraper | Public contact emails from Unreal Engine Marketplace |
| WordPress Plugin Directory Email Scraper | Public contact emails from WordPress Plugin Directory |
| WordPress Theme Directory Email Scraper | Public contact emails from WordPress Theme Directory |
| Zapier App Directory Email Scraper | Public contact emails from Zapier App Directory |
| App Store Email and Phone Number Scraper | Emails and phone numbers from App Store |
| Atlassian Marketplace Email and Phone Number Scraper | Emails and phone numbers from Atlassian Marketplace |
| Bitbucket Email and Phone Number Scraper | Emails and phone numbers from Bitbucket |
| Chrome Web Store Email and Phone Number Scraper | Emails and phone numbers from Chrome Web Store |
| CodePen Email and Phone Number Scraper | Emails and phone numbers from CodePen |
| Confluence Email and Phone Number Scraper | Emails and phone numbers from Confluence |
| DEV Community Email and Phone Number Scraper | Emails and phone numbers from DEV Community |
| Docker Hub Email and Phone Number Scraper | Emails and phone numbers from Docker Hub |
| Figma Community Email and Phone Number Scraper | Emails and phone numbers from Figma Community |
| Firefox Add-ons Email and Phone Number Scraper | Emails and phone numbers from Firefox Add-ons |
| GitHub Email and Phone Number Scraper | Emails and phone numbers from GitHub |
| GitLab Email and Phone Number Scraper | Emails and phone numbers from GitLab |
Leave a review
If the SourceForge Email Scraper saved you time, please leave a star rating and a short review on the Actor page.
Reviews are how other buyers judge whether a tool works, and they tell us which features to build next.
If something did not work, email neurodata.apify@gmail.com instead - bugs get fixed faster than they get complained about.
Support
Questions, bug reports, or a custom build of the SourceForge Email Scraper tuned to your own domain list? Email neurodata.apify@gmail.com.