# SourceForge Email Scraper (`leads-scraper/sourceforge-email-scraper`) Actor

SourceForge Email Scraper SD - SourceForge Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from SourceForge results by keyword, location and email domain - SourceForge email extractor.

- **URL**: https://apify.com/leads-scraper/sourceforge-email-scraper.md
- **Developed by:** [Leads Scraper](https://apify.com/leads-scraper) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.49 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### SourceForge Email Scraper - maintainer contacts from project and user pages

The **SourceForge Email Scraper** is an Apify Actor that collects publicly indexed contact emails from SourceForge project pages, user profiles, wikis and ticket threads. Give it keywords, an optional location and your email domains, and it returns a structured lead dataset.

SourceForge is the long memory of open source. Projects that predate GitHub are still hosted, still downloaded and still maintained there, and many of them publish a maintainer address in plain text on the project page.

That is exactly what the SourceForge Email Scraper is for: reaching the people behind mature, long-lived open-source software that never migrated anywhere else.

#### What the SourceForge Email Scraper reads

It reads **only Google's public index** of `sourceforge.net`. It builds `site:` queries, fetches result pages through the Apify GOOGLE_SERP proxy, and extracts emails from the result titles and snippets.

It does **not** read git or Subversion history, clone repositories, or use any SourceForge API. There is no login, no browser, no JavaScript rendering and no cookies in the pipeline.

Addresses surface from project summary pages, `README` and documentation files, wiki pages, ticket and discussion threads, and user profile pages under `/u/` - wherever a maintainer chose to publish one and Google indexed it.

#### Who the SourceForge Email Scraper is built for

Security researchers contacting maintainers of dependencies, developer relations teams reaching legacy-project owners, agencies doing developer lead generation in niche verticals, and anyone who needs to talk to the person still shipping a twenty-year-old utility.

If you have hunted through a SourceForge project page for a way to email its owner, this Actor does that at scale and returns a spreadsheet.

***

### Key features of the SourceForge Email Scraper

| Feature | What it does |
|---|---|
| Google `site:` dorking | Queries `sourceforge.net` through the Apify GOOGLE_SERP proxy |
| Query expansion | Base, quoted and `intitle:` variants plus one variant per query modifier; base queries run first |
| Domain-filtered extraction | Keeps only emails ending in your `customDomains` values |
| Global deduplication | One row per unique email address across every query and page |
| Obfuscation handling | Understands `name [at] domain [dot] com`, `name (at) domain`, `name @ domain.com`, `domain .com`, zero-width characters and the full-width `＠` - common in older project pages written to dodge spam bots |
| Junk filter | Rejects placeholders such as `email@`, `yourname@`, `test@`, `xxx@` and single-character locals |
| Boundary-correct matching | `@gmail.com` never matches inside `@gmail.company` or `@gmail.com.br` |
| Soft-wrap repair | Discards a hit that is only the tail of another email in the same result block |
| Structural parsing | Locates the `<h3>` title then the smallest surrounding block, not Google's CSS class names |
| Whole-page fallback | A Google markup change degrades the run to "emails without account details", not "no emails" |
| Block detection | CAPTCHA, "unusual traffic" and consent pages are detected and retried rather than treated as empty |
| Retries and backoff | Up to 3 attempts per page, exponential backoff, a fresh proxy session per request |
| Requeue of failures | Blocked or failed queries are re-queued once at the end of the run |
| Async concurrency | An `asyncio` worker pool with a shared stop signal on `maxEmails` |
| Resumable state | Key-value-store progress keyed by an input hash, with throttled saves plus saves on `PERSIST_STATE`, `MIGRATING` and `ABORTING` |
| Streaming output | Leads reach the dataset as they are found, so an aborted run keeps what it collected |
| Run summary | Logs pages fetched, blocked pages, retries and emails per page |

The obfuscation handling is unusually relevant on SourceForge. Older project pages routinely write `maintainer [at] example [dot] com`, and the SourceForge Email Scraper normalises those into real addresses.

***

### How the SourceForge Email Scraper works

The SourceForge Email Scraper pipeline is six steps, with no authentication anywhere in it.

1. It reads your input: keywords, location, email domains and limits.
2. It builds Google queries with the `site:` operator, for example `site:sourceforge.net maintainer contact "@gmail.com"`.
3. It fetches Google result pages asynchronously with `aiohttp` through the Apify GOOGLE_SERP proxy.
4. It parses each result block structurally, finding the `<h3>` title and the smallest block around it.
5. It extracts email addresses from the block text using a domain-filtered regex.
6. It deduplicates globally and pushes each lead straight into the Apify dataset.

#### Why query expansion matters more here

SourceForge is a single domain, so there is no second site to spread queries across the way `github.io` complements `github.com`.

Google caps one query at roughly 300 results. With only one domain to target, query expansion is the main lever the SourceForge Email Scraper has for increasing total coverage.

The default `queryModifiers` - `email`, `contact`, `maintainer`, `author`, `support` - match how SourceForge project pages label their contact sections.

#### What the SourceForge Email Scraper does not do

No SourceForge login, no API, no repository or SVN checkout, no commit-history parsing, no JavaScript rendering. It is an independent Apify Actor, not affiliated with or endorsed by SourceForge or Slashdot Media.

***

### SourceForge Email Scraper input fields

Only `keywords` is required by the SourceForge Email Scraper. All defaults below come from the Actor's shipped input schema.

| Field | Type | Default | Meaning |
|---|---|---|---|
| `keywords` | array (required) | `["maintainer", "open source"]` | Search terms describing the SourceForge accounts and projects you want (niche, job title, technology) |
| `location` | string | `""` | Optional location phrase added to every query |
| `customDomains` | array | `["@gmail.com", "@yahoo.com"]` | Only emails on these domains are kept; the leading `@` is optional |
| `maxEmails` | integer 1-10000 | `20` | Stop after this many unique emails |
| `countryCode` | string | `""` | Two-letter country for the search proxy (US, GB, DE...) |
| `expandQueries` | boolean | `true` | Search each keyword x domain pair in several phrasings |
| `queryModifiers` | array | `["email", "contact", "maintainer", "author", "support"]` | Extra words combined with each keyword when expansion is on |
| `maxPagesPerQuery` | integer 1-50 | `30` | Page cap per query |
| `maxConcurrency` | integer 1-20 | `5` | Parallel queries |

#### Example SourceForge Email Scraper input

```json
{
  "keywords": ["cnc controller", "scientific computing", "embedded toolchain"],
  "location": "",
  "customDomains": ["@gmail.com", "@yahoo.com"],
  "maxEmails": 300,
  "countryCode": "US",
  "expandQueries": true,
  "queryModifiers": ["email", "contact", "maintainer", "author", "support"],
  "maxPagesPerQuery": 30,
  "maxConcurrency": 5
}
```

#### Tuning the SourceForge Email Scraper

Domain-specific project vocabulary works best in the SourceForge Email Scraper. `plc simulator`, `astronomy imaging`, `serial terminal` and similar phrases match how SourceForge projects actually describe themselves.

Because there is only one site to query, add email domains generously - `@yahoo.com`, `@hotmail.com` and older free providers are well represented in this population.

Keep `expandQueries` on. With a single domain to target, disabling it is the fastest way to cap the SourceForge Email Scraper at a few hundred results.

***

### SourceForge Email Scraper output fields

Every dataset item the SourceForge Email Scraper writes has all 14 fields. Values are empty or `null` when SourceForge did not expose them in Google's result - they are never dropped.

| Field | Meaning |
|---|---|
| `network` | Platform name |
| `keyword` | The keyword that produced the lead |
| `query` | The exact Google query used |
| `title` | Raw result title |
| `accountName` | Account label Google prints - a user handle on profile pages, a project name on project pages |
| `fullName` | Display name parsed from a profile-style title; empty for project and doc pages |
| `username` | URL-safe SourceForge handle when one is exposed; otherwise `null` |
| `profileUrl` | Canonical `https://sourceforge.net/u/{username}/profile/` URL when a handle is known; otherwise empty |
| `url` | Direct platform link when exposed, else the profile URL |
| `description` | Snippet text, cleaned of labels and engagement counters |
| `email` | Lower-cased email address |
| `emailDomain` | The matched domain, for example `@gmail.com` |
| `possiblyTruncated` | `true` when Google's snippet ellipsis touched the email - verify before sending |
| `foundAt` | ISO 8601 UTC timestamp |

#### Example SourceForge Email Scraper output

```json
[
  {
    "network": "SourceForge",
    "keyword": "cnc controller",
    "query": "site:sourceforge.net cnc controller maintainer \"@gmail.com\"",
    "title": "mgrieder - SourceForge",
    "accountName": "mgrieder",
    "fullName": "Martin Grieder",
    "username": "mgrieder",
    "profileUrl": "https://sourceforge.net/u/mgrieder/profile/",
    "url": "https://sourceforge.net/u/mgrieder/profile/",
    "description": "Maintainer of a stepper motion control suite. Patches and questions to martin.grieder.cnc@gmail.com",
    "email": "martin.grieder.cnc@gmail.com",
    "emailDomain": "@gmail.com",
    "possiblyTruncated": false,
    "foundAt": "2026-08-31T13:10:33.402Z"
  },
  {
    "network": "SourceForge",
    "keyword": "scientific computing",
    "query": "site:sourceforge.net scientific computing contact \"@yahoo.com\"",
    "title": "AstroReduce - Browse Files at SourceForge.net",
    "accountName": "AstroReduce",
    "fullName": "",
    "username": null,
    "profileUrl": "",
    "url": "https://sourceforge.net/projects/astroreduce/files/",
    "description": "Bug reports to astroreduce.project@yahoo.com. Release notes for 4.2 are in the wiki ...",
    "email": "astroreduce.project@yahoo.com",
    "emailDomain": "@yahoo.com",
    "possiblyTruncated": true,
    "foundAt": "2026-08-31T13:10:51.860Z"
  },
  {
    "network": "SourceForge",
    "keyword": "embedded toolchain",
    "query": "site:sourceforge.net intitle:\"embedded toolchain\" support \"@gmail.com\"",
    "title": "Ticket #148: cross compiler build fails on ARM",
    "accountName": "",
    "fullName": "",
    "username": null,
    "profileUrl": "",
    "url": "https://sourceforge.net/p/xtool-embedded/tickets/148/",
    "description": "Reported by the toolchain author, reachable at xtool.embedded.dev@gmail.com for build issues",
    "email": "xtool.embedded.dev@gmail.com",
    "emailDomain": "@gmail.com",
    "possiblyTruncated": false,
    "foundAt": "2026-08-31T13:11:24.517Z"
  }
]
```

This is a representative mix of SourceForge Email Scraper output. The first row is a user profile with a full identity, the second a project page with a project name but no handle, the third a ticket thread with an email and no account details at all.

***

### Use cases for the SourceForge Email Scraper

| Use case | How the SourceForge Email Scraper helps |
|---|---|
| Security disclosure | Reach maintainers of legacy dependencies who have no `SECURITY.md` and no issue tracker you can use |
| Open-source outreach | Contact project owners about forks, ports, packaging or maintenance handover |
| Developer relations | Find owners of long-lived projects in your category before a launch or migration campaign |
| Developer lead generation | Build a niche list of technical project owners in vertical software |
| Vendor and partner discovery | Identify who still maintains tooling your customers depend on |
| Migration and modernisation offers | Approach owners of projects that would benefit from a newer platform |
| Academic and research outreach | Reach authors of scientific and engineering utilities published on SourceForge |
| Ecosystem and archival research | Measure how many projects in a niche still publish a working contact address |

#### The security disclosure angle

Plenty of software in production dependency trees still lives on SourceForge, and much of it has no modern reporting channel. The SourceForge Email Scraper often finds the only working contact route.

Search your dependency names as keywords, set the domains you expect, and check `possiblyTruncated` before you send anything sensitive.

#### Combining with sibling Actors

Run the SourceForge Email Scraper together with the [GitHub Email Scraper](https://apify.com/leads-scraper/github-email-scraper) and the **PyPI Email Scraper** to catch maintainers who publish source in one place and packages in another.

***

### Expected results from the SourceForge Email Scraper

In live test runs against Google, about **4 out of 10 parsed results carried an account identity** - a handle, a display name, or both. Output is a genuine mix of user pages and project pages.

User profile pages under `/u/` resolve cleanly to a handle and a profile URL. Project pages, download listings, wikis and tickets typically give you a project name instead, and sometimes nothing but the email itself.

Every SourceForge Email Scraper row still carries `email`, `emailDomain`, `url`, `title`, `description` and `query`, which is enough to qualify a lead. Volume depends on keywords and domains; no yield is guaranteed.

***

### Limitations of the SourceForge Email Scraper

- **Publicly indexed emails only.** An address is findable only when it is already visible in Google's index. Anything behind a login or never published is unreachable.
- **No repository history, no API.** The SourceForge Email Scraper does not read git or SVN history, clone repositories, or call any SourceForge API. Only Google result titles and snippets are parsed.
- **Google's ~300-result cap.** One query returns roughly 300 results at most. Because SourceForge is a single domain, `expandQueries` is the main way to go beyond that.
- **`possiblyTruncated`.** When `true`, Google's snippet ellipsis touched the email and it may be cut off. Verify those rows before sending.
- **`username` and `profileUrl` are often empty.** They are populated only when SourceForge exposes a handle in the result. Project, download, wiki and ticket pages usually show a project name instead, leaving `accountName` filled but `username` `null`.
- **Stale contacts.** Many SourceForge projects are old. An address published a decade ago may no longer be monitored - treat bounces as expected.
- **Apify GOOGLE_SERP proxy required.** The Actor cannot run without Apify proxy credentials.
- **Free plan cap.** Free Apify plans are limited to 100 emails per run; paid plans are uncapped.
- **Variable results.** Output depends on keywords, domains and location, and Google's index changes over time.

***

### Responsible use

Maintainer emails the SourceForge Email Scraper returns were published for project correspondence - bug reports, patches, packaging questions, security disclosures. That is the spirit in which to use them.

If you contact people in the EU or UK, GDPR applies: have a lawful basis, identify yourself, say where you found the address, and honour opt-outs immediately.

Many of these projects are maintained by one person in their spare time. Be specific, be brief, and do not send a maintainer a marketing sequence.

***

### SourceForge Email Scraper FAQ

#### Does the SourceForge Email Scraper read repository history or use an API?

No. It does not read git or SVN history, clone repositories, or call any SourceForge API. It parses Google search results for already-indexed public pages on `sourceforge.net`.

#### Do I need a SourceForge account?

No. There is no authentication, no browser, no JavaScript rendering and no cookies. The only credential involved is your Apify account's GOOGLE_SERP proxy access.

#### Where does the SourceForge Email Scraper get the emails?

From Google's result titles and snippets: project summary pages, `README` and documentation files, wiki pages, ticket and discussion threads, and user profile pages under `/u/`.

#### How often do SourceForge Email Scraper rows include a profile URL?

In live testing, about 4 in 10 parsed results carried an account identity. User profile pages resolve well; project, download and ticket pages usually do not.

#### Why is `username` sometimes `null`?

Because Google exposed a project name rather than a SourceForge handle for that result. The row is still a usable lead with `email`, `accountName`, `url` and `description`.

#### What does `possiblyTruncated: true` mean?

Google's snippet ellipsis touched the address, so it may be incomplete. Verify those rows before you use them.

#### Can the SourceForge Email Scraper filter to specific domains?

Yes. Set `customDomains` to what you want, for example `["@gmail.com", "@yahoo.com", "@acme.io"]`. The leading `@` is optional and only matching addresses are kept.

#### Does the SourceForge Email Scraper decode `maintainer [at] example [dot] com`?

Yes. That style is common on older SourceForge pages, and the extractor normalises `[at]`, `(at)`, spaced `@`, spaced `.com`, zero-width characters and the full-width `＠`.

#### How many emails can one SourceForge Email Scraper run return?

`maxEmails` accepts 1 to 10000 and defaults to 20. Free Apify plans are capped at 100 emails per run; paid plans are uncapped.

#### How do I target one country?

Set `countryCode` to a two-letter code such as `US`, `GB` or `DE`, and put a city or region in `location` so it is appended to every query.

#### What happens if a SourceForge Email Scraper run is interrupted or migrated?

Progress is stored in the key-value store keyed by a hash of your input, with saves on `PERSIST_STATE`, `MIGRATING` and `ABORTING`. Leads already pushed to the dataset are kept.

#### Is the SourceForge Email Scraper affiliated with SourceForge?

No. It is an independent Apify Actor, not supported, endorsed or affiliated with SourceForge or Slashdot Media.

***

### Related Actors

The SourceForge Email Scraper belongs to a Developer & Technology family of Apify Actors that apply the same method to different platforms. Run several to cover an ecosystem instead of one site.

| Actor | What it collects |
|---|---|
| [SourceForge Email and Phone Number Scraper](https://apify.com/neuro-scraper/sourceforge-email-and-phone-number-scraper) | Emails and phone numbers from SourceForge |
| [SourceForge Phone Number Scraper](https://apify.com/neuro-scraper/sourceforge-phone-number-scraper) | Public phone numbers from SourceForge |
| [App Store Email Scraper](https://apify.com/leads-scraper/app-store-email-scraper) | Public contact emails from App Store |
| [Atlassian Marketplace Email Scraper](https://apify.com/leads-scraper/atlassian-marketplace-email-scraper) | Public contact emails from Atlassian Marketplace |
| [Bitbucket Email Scraper](https://apify.com/neuro-scraper/bitbucket-email-scraper) | Public contact emails from Bitbucket |
| [Chrome Web Store Email Scraper](https://apify.com/leads-scraper/chrome-web-store-email-scraper) | Public contact emails from Chrome Web Store |
| [CodePen Email Scraper](https://apify.com/neuro-scraper/codepen-email-scraper) | Public contact emails from CodePen |
| [Confluence Email Scraper](https://apify.com/neuro-scraper/confluence-email-scraper) | Public contact emails from Confluence |
| [Dev.to Email Scraper](https://apify.com/neuro-scraper/dev-to-email-scraper) | Public contact emails from DEV Community |
| [Docker Hub Email Scraper](https://apify.com/neuro-scraper/docker-hub-email-scraper) | Public contact emails from Docker Hub |
| [Figma Community Email Scraper](https://apify.com/neuro-scraper/figma-community-email-scraper) | Public contact emails from Figma Community |
| [Firefox Add-ons Email Scraper](https://apify.com/neuro-scraper/firefox-add-ons-email-scraper) | Public contact emails from Firefox Add-ons |
| [GitHub Email Scraper](https://apify.com/leads-scraper/github-email-scraper) | Public contact emails from GitHub |
| [GitLab Email Scraper](https://apify.com/neuro-scraper/gitlab-email-scraper) | Public contact emails from GitLab |
| [Google Play Email Scraper](https://apify.com/leads-scraper/google-play-email-scraper) | Public contact emails from Google Play |
| [Hashnode Email Scraper](https://apify.com/neuro-scraper/hashnode-email-scraper) | Public contact emails from Hashnode |
| [HubSpot Marketplace Email Scraper](https://apify.com/leads-scraper/hubspot-marketplace-email-scraper) | Public contact emails from HubSpot Marketplace |
| [Hugging Face Email Scraper](https://apify.com/leads-scraper/hugging-face-email-scraper) | Public contact emails from Hugging Face |
| [Jira Email Scraper](https://apify.com/neuro-scraper/jira-email-scraper) | Public contact emails from Jira |
| [Maven Central Email Scraper](https://apify.com/neuro-scraper/maven-central-email-scraper) | Public contact emails from Maven Central |
| [Microsoft AppSource Email Scraper](https://apify.com/neuro-scraper/microsoft-appsource-email-scraper) | Public contact emails from Microsoft AppSource |
| [Salesforce AppExchange Email Scraper](https://apify.com/leads-scraper/salesforce-appexchange-email-scraper) | Public contact emails from Salesforce AppExchange |
| [Shopify App Store Email Scraper](https://apify.com/leads-scraper/shopify-app-store-email-scraper) | Public contact emails from Shopify App Store |
| [Slack App Directory Email Scraper](https://apify.com/leads-scraper/slack-app-directory-email-scraper) | Public contact emails from Slack App Directory |
| [Stack Overflow Email Scraper](https://apify.com/leads-scraper/stack-overflow-email-scraper) | Public contact emails from Stack Overflow |
| [Unity Asset Store Email Scraper](https://apify.com/leads-scraper/unity-asset-store-email-scraper) | Public contact emails from Unity Asset Store |
| [Unreal Engine Marketplace Email Scraper](https://apify.com/leads-scraper/unreal-engine-marketplace-email-scraper) | Public contact emails from Unreal Engine Marketplace |
| [WordPress Plugin Directory Email Scraper](https://apify.com/neuro-scraper/wordpress-plugin-directory-email-scraper) | Public contact emails from WordPress Plugin Directory |
| [WordPress Theme Directory Email Scraper](https://apify.com/leads-scraper/wordpress-theme-directory-email-scraper) | Public contact emails from WordPress Theme Directory |
| [Zapier App Directory Email Scraper](https://apify.com/leads-scraper/zapier-app-directory-email-scraper) | Public contact emails from Zapier App Directory |
| [App Store Email and Phone Number Scraper](https://apify.com/leads-scraper/app-store-email-and-phone-number-scraper) | Emails and phone numbers from App Store |
| [Atlassian Marketplace Email and Phone Number Scraper](https://apify.com/neuro-scraper/atlassian-marketplace-email-and-phone-number-scraper) | Emails and phone numbers from Atlassian Marketplace |
| [Bitbucket Email and Phone Number Scraper](https://apify.com/neuro-scraper/bitbucket-email-and-phone-number-scraper) | Emails and phone numbers from Bitbucket |
| [Chrome Web Store Email and Phone Number Scraper](https://apify.com/neuro-scraper/chrome-web-store-email-and-phone-number-scraper) | Emails and phone numbers from Chrome Web Store |
| [CodePen Email and Phone Number Scraper](https://apify.com/neuro-scraper/codepen-email-and-phone-number-scraper) | Emails and phone numbers from CodePen |
| [Confluence Email and Phone Number Scraper](https://apify.com/neuro-scraper/confluence-email-and-phone-number-scraper) | Emails and phone numbers from Confluence |
| [DEV Community Email and Phone Number Scraper](https://apify.com/neuro-scraper/dev-to-email-and-phone-number-scraper) | Emails and phone numbers from DEV Community |
| [Docker Hub Email and Phone Number Scraper](https://apify.com/neuro-scraper/docker-hub-email-and-phone-number-scraper) | Emails and phone numbers from Docker Hub |
| [Figma Community Email and Phone Number Scraper](https://apify.com/neuro-scraper/figma-community-email-and-phone-number-scraper) | Emails and phone numbers from Figma Community |
| [Firefox Add-ons Email and Phone Number Scraper](https://apify.com/neuro-scraper/firefox-add-ons-email-and-phone-number-scraper) | Emails and phone numbers from Firefox Add-ons |
| [GitHub Email and Phone Number Scraper](https://apify.com/leads-scraper/github-email-and-phone-number-scraper) | Emails and phone numbers from GitHub |
| [GitLab Email and Phone Number Scraper](https://apify.com/neuro-scraper/gitlab-email-and-phone-number-scraper) | Emails and phone numbers from GitLab |

### Leave a review

If the SourceForge Email Scraper saved you time, please leave a star rating and a short review
on the Actor page.

Reviews are how other buyers judge whether a tool works, and they tell us which features to
build next.

If something did not work, email <neurodata.apify@gmail.com>
instead - bugs get fixed faster than they get complained about.

### Support

Questions, bug reports, or a custom build of the SourceForge Email Scraper tuned to your own domain list? Email **neurodata.apify@gmail.com**.

# Actor input Schema

## `keywords` (type: `array`):

Search terms describing the SourceForge accounts you want (niche, job title, industry).

## `location` (type: `string`):

Optional location phrase added to every query (e.g. "New York").

## `customDomains` (type: `array`):

Only emails ending with one of these domains are collected. With or without the leading @. Each domain is searched separately, so more domains means more results but a longer run - remove some for a faster, narrower search, or add your own (e.g. @company.com).

## `maxEmails` (type: `integer`):

Stop once this many unique emails have been collected.

## `countryCode` (type: `string`):

Two-letter country code for the search proxy (e.g. US, GB, DE). Empty for any.

## `expandQueries` (type: `boolean`):

Search each keyword x domain pair with several phrasings. Recommended - Google caps a single query at ~300 results.

## `queryModifiers` (type: `array`):

Extra words combined with each keyword when Expand queries is on. Tuned for SourceForge.

## `maxPagesPerQuery` (type: `integer`):

Google rarely returns more than ~30 pages for one query.

## `maxConcurrency` (type: `integer`):

How many queries run in parallel.

## Actor input object example

```json
{
  "keywords": [
    "maintainer",
    "open source"
  ],
  "location": "",
  "customDomains": [
    "@gmail.com",
    "@yahoo.com"
  ],
  "maxEmails": 20,
  "countryCode": "",
  "expandQueries": true,
  "queryModifiers": [
    "email",
    "contact",
    "maintainer",
    "author",
    "support"
  ],
  "maxPagesPerQuery": 30,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

Records produced by SourceForge Email Scraper, stored in the run's default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "maintainer",
        "open source"
    ],
    "location": "",
    "customDomains": [
        "@gmail.com",
        "@yahoo.com"
    ],
    "countryCode": "",
    "queryModifiers": [
        "email",
        "contact",
        "maintainer",
        "author",
        "support"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("leads-scraper/sourceforge-email-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": [
        "maintainer",
        "open source",
    ],
    "location": "",
    "customDomains": [
        "@gmail.com",
        "@yahoo.com",
    ],
    "countryCode": "",
    "queryModifiers": [
        "email",
        "contact",
        "maintainer",
        "author",
        "support",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("leads-scraper/sourceforge-email-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "maintainer",
    "open source"
  ],
  "location": "",
  "customDomains": [
    "@gmail.com",
    "@yahoo.com"
  ],
  "countryCode": "",
  "queryModifiers": [
    "email",
    "contact",
    "maintainer",
    "author",
    "support"
  ]
}' |
apify call leads-scraper/sourceforge-email-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,leads-scraper/sourceforge-email-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/VeLayh6cUkd3o98ap/builds/WlwjpNBxOZcKB4dOb/openapi.json
