# SimilarWeb Email Scraper (`email_scraper/similarweb-email-scraper`) Actor

SimilarWeb Email Scraper extracts publicly indexed email addresses using targeted keywords, location filters, custom email domains, and exclusion terms. Build structured SimilarWeb contact datasets for business research, lead discovery, company research, and contact analysis.

- **URL**: https://apify.com/email\_scraper/similarweb-email-scraper.md
- **Developed by:** [Email Scraper](https://apify.com/email_scraper) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.49 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### SimilarWeb Email Scraper

**SimilarWeb Email Scraper** is an Apify Actor for discovering publicly indexed email addresses associated with SimilarWeb pages. It uses targeted keywords, optional geographic terms, configurable email-domain suffixes, and exclusion filters to identify relevant SimilarWeb search results and return structured contact records.

You provide one or more search keywords, optionally specify a country, state, or city, choose the email domains you want to target, and set an email collection limit for each keyword + domain combination. The Actor returns structured records containing the keyword, result title, description, URL, and discovered email address.

This makes the SimilarWeb Email Scraper useful for business research, contact discovery, market research, company research, lead research, and structured dataset creation based on publicly indexed information.

### What Is a SimilarWeb Email Scraper?

A SimilarWeb Email Scraper is designed to find email addresses appearing in publicly indexed SimilarWeb search-result descriptions.

Instead of manually searching SimilarWeb-related pages for contact information, you can provide targeted search terms and let the Actor process multiple keyword and email-domain combinations.

The Actor supports optional location targeting, custom email-domain suffixes, and exclusion words or phrases. Results are delivered as structured dataset records that can be reviewed and used for research or other data workflows.

The quality and quantity of results depend on what is publicly indexed and available for the selected searches. The Actor does not guarantee that every requested email target will be available.

### Key Features

| Feature               | Description                                                            | User Benefit                                     |
| --------------------- | ---------------------------------------------------------------------- | ------------------------------------------------ |
| Keyword-based search  | Search using one or multiple keywords or queries                       | Target specific business or professional topics  |
| Location filtering    | Optionally add a country, state, or city                               | Narrow searches geographically                   |
| Custom email domains  | Specify domains such as `@gmail.com` or `@yahoo.com`                   | Focus collection on relevant email types         |
| Per-combination limit | Set the maximum number of emails for each keyword + domain combination | Control the depth of each search                 |
| Exclude words         | Skip descriptions containing selected words or phrases                 | Reduce unwanted results                          |
| Duplicate prevention  | Previously collected email addresses are not added again               | Keep the resulting dataset cleaner               |
| Structured dataset    | Results contain consistent fields                                      | Make collected data easier to review and analyze |
| Multiple keywords     | Process several search terms in one run                                | Build broader search coverage                    |

### What Data Can You Extract?

The SimilarWeb Email Scraper returns five user-facing data fields for each collected email.

- **Keyword** — The search keyword associated with the result.
- **Title** — The title of the matching search result.
- **Description** — The publicly indexed description or snippet associated with the result.
- **URL** — The URL associated with the search result.
- **Email** — The email address identified in the result description that matches one of the configured domain suffixes.

These fields provide both the discovered contact information and the surrounding search-result context, making it easier to understand where each email was found.

The dataset is therefore more than a simple list of email addresses. It also preserves the keyword, title, description, and URL associated with each result.

### Why Use This Actor?

Manual contact research can require repeatedly testing different search terms, checking result descriptions, identifying matching email domains, and recording the information in a structured format.

The SimilarWeb Email Scraper automates this repetitive collection workflow.

You can define several related keywords instead of relying on one broad search term. For example, a business research workflow might use terms such as `marketing manager`, `digital marketing`, `online business`, or `business consultant`.

You can also combine those keywords with different email-domain suffixes and optional geographic terms. This gives you a configurable way to organize searches around the information you actually need.

### Benefits

- **Automated email discovery** — Reduce repetitive manual searching for publicly indexed email information.
- **Structured results** — Receive consistent keyword, title, description, URL, and email fields.
- **Flexible targeting** — Combine multiple search keywords with selected email domains.
- **Geographic refinement** — Add a country, state, or city when location-specific research is needed.
- **Result filtering** — Exclude descriptions containing unwanted words or phrases.
- **Research-friendly output** — Keep contextual information alongside each discovered email.
- **Scalable search configuration** — Configure several keyword and domain combinations within one run.
- **Duplicate control** — The Actor avoids adding the same email address repeatedly during a run.

### How to Use the SimilarWeb Email Scraper

The basic workflow is straightforward:

1. Enter one or more keywords in the `keywords` field.
2. Optionally enter a country, state, or city in `location`.
3. Add the email-domain suffixes you want to search for in `customDomains`.
4. Set `maxEmails` according to the desired collection depth.
5. Optionally add unwanted words or phrases to `excludeWords`.
6. Start the Actor.
7. Review the resulting dataset containing the matching search context and email addresses.

For broader research, use several specific keywords rather than one very broad term. More targeted queries can help distinguish different types of businesses, professionals, or topics.

### Input

The Actor requires the `keywords` field. All other configuration fields are optional.

| Field           | Type             | Required | Default                                    | Description                                                                             |
| --------------- | ---------------- | -------- | ------------------------------------------ | --------------------------------------------------------------------------------------- |
| `keywords`      | Array of strings | Yes      | `["marketing manager", "online business"]` | Search keywords or queries used to identify relevant SimilarWeb results.                |
| `location`      | String           | No       | `""`                                       | Optional country, state, or city used to narrow the search.                             |
| `customDomains` | Array of strings | No       | `["@gmail.com"]`                           | Email-domain suffixes to target, such as `@gmail.com`, `@yahoo.com`, or `@outlook.com`. |
| `maxEmails`     | Integer          | No       | `5`                                        | Maximum target for each keyword + domain combination. Allowed range is 1–10,000.        |
| `excludeWords`  | Array of strings | No       | `[]`                                       | Words or phrases that cause matching result descriptions to be skipped.                 |

The `keywords` field uses a list of strings, allowing multiple search terms in the same run.

The `location` field can be left empty when geographic filtering is not required.

For `customDomains`, enter the complete email suffix including the `@` symbol. You can specify multiple domains.

The `maxEmails` value applies independently to each keyword + domain combination. For example, two keywords and three domains create six combinations, with the configured target applied to each combination.

For free users, the Actor limits the effective `maxEmails` configuration to 100 when a higher value or no value is supplied. The Actor's code applies this ceiling to the configured per-combination limit.

### Input Example

```json
{
  "keywords": [
    "marketing manager",
    "digital marketing",
    "online business"
  ],
  "location": "United States",
  "customDomains": [
    "@gmail.com",
    "@yahoo.com",
    "@outlook.com"
  ],
  "maxEmails": 20,
  "excludeWords": [
    "crypto",
    "onlyfans"
  ]
}
```

### Output

Each collected result is added to the Apify dataset as a structured record.

| Field         | Description                                                                                     |
| ------------- | ----------------------------------------------------------------------------------------------- |
| `keyword`     | The keyword used for the search that produced the result.                                       |
| `title`       | Title of the matching search result.                                                            |
| `description` | Description or search-result snippet associated with the result.                                |
| `url`         | URL associated with the matching result.                                                        |
| `email`       | Email address extracted from the result description that matches the configured domain pattern. |

The Actor can produce multiple records from different keyword and domain combinations. Duplicate email addresses are not added again when they have already been collected during the run.

The output provides context around each email, which can be useful when reviewing or validating collected contact records.

### Output Example

```json
{
  "keyword": "marketing manager",
  "title": "Example Marketing Company",
  "description": "Marketing services and business information. Contact: example@gmail.com",
  "url": "https://www.similarweb.com/example",
  "email": "example@gmail.com"
}
```

The example above demonstrates the output structure. Actual titles, descriptions, URLs, and email addresses depend on the publicly indexed search results available for the configured inputs.

### Use Cases

The SimilarWeb Email Scraper can support several research and data-collection workflows.

- **Business research** — Find publicly indexed contact information associated with relevant SimilarWeb results.
- **Lead research** — Build structured contact datasets around selected professional or business keywords.
- **Market research** — Investigate businesses or topics using multiple targeted search terms.
- **Company research** — Organize search results and contact information around specific business categories.
- **Industry research** — Search across professional terminology, industries, or business roles.
- **Geographic research** — Combine keywords with locations to focus searches on particular markets.
- **Dataset creation** — Build structured datasets containing email addresses and their associated search context.
- **Data analysis** — Use keyword, description, URL, and email fields for subsequent research and analysis.
- **Contact discovery** — Identify publicly indexed email addresses matching selected domain suffixes.

### Competitive Advantages

The Actor provides several practical configuration options in a single workflow.

- Multiple keywords can be processed in one run.
- Multiple email-domain suffixes can be configured.
- Geographic filtering is available when location matters.
- Exclusion words and phrases can filter unwanted descriptions.
- The resulting dataset preserves search context instead of returning only email addresses.
- Each keyword + domain combination has its own configured collection target.
- Duplicate email addresses are prevented from being added repeatedly.

These capabilities make the Actor adaptable to different research scopes without requiring users to manually repeat the same search process.

### Advantages

- Clear input structure with one required field and several optional controls.
- Supports targeted rather than single-query research.
- Produces structured Apify dataset records.
- Includes contextual search-result information with each email.
- Supports custom email-domain targeting.
- Provides configurable exclusion filtering.
- Supports geographic search refinement.
- Allows users to control collection depth through `maxEmails`.

### Limitations

The Actor works with publicly indexed search-result information, so results depend on what is available for the selected keywords, domains, and optional location.

Important limitations include:

- Not every SimilarWeb page or business will have a publicly indexed email address.
- A requested `maxEmails` value does not guarantee that the target number will be found.
- Narrow keywords or locations may produce fewer results.
- The selected email-domain suffixes determine which email addresses can be collected.
- Exclusion terms can intentionally remove otherwise matching results.
- Search availability can change over time.
- The Actor does not guarantee that every returned email address remains active or belongs to the intended organization.
- Free users have an effective `maxEmails` ceiling of 100 when a higher or unspecified value is supplied.

Because the Actor depends on publicly indexed information, empty or partial results should be treated as a possible outcome rather than an indication that a specific number of contacts must exist.

### Pros and Cons

| Pros                          | Cons                                                                       |
| ----------------------------- | -------------------------------------------------------------------------- |
| Multiple keyword support      | Results depend on publicly indexed information                             |
| Custom email-domain targeting | Narrow searches may return few emails                                      |
| Optional geographic filtering | Requested limits do not guarantee matching results                         |
| Exclusion words and phrases   | Collected email addresses should be independently validated when important |
| Structured contextual output  | Search coverage can vary over time                                         |
| Duplicate email prevention    | Only configured email-domain patterns are targeted                         |

### Comparison With Alternative Approaches

| Capability             | SimilarWeb Email Scraper                      | Manual / Typical Alternative    |
| ---------------------- | --------------------------------------------- | ------------------------------- |
| Keyword-based research | Supported                                     | Usually performed manually      |
| Multiple search terms  | Supported                                     | Requires repeated searches      |
| Email-domain filtering | Configurable                                  | Often requires manual filtering |
| Location refinement    | Supported                                     | Depends on the search workflow  |
| Exclusion filtering    | Supported                                     | Usually performed manually      |
| Structured output      | Dataset fields for each result                | May require manual formatting   |
| Duplicate prevention   | Supported during collection                   | Often requires separate cleanup |
| Search-result context  | Title, description, URL, and keyword included | May need manual recording       |

This comparison describes workflow differences rather than claiming that one approach is universally better.

### Best Practices

- Start with a small `maxEmails` value to test the search configuration.
- Use several specific keywords instead of relying only on a broad keyword.
- Add related job titles, business categories, or professional terms when appropriate.
- Use multiple email domains when you want broader email coverage.
- Leave `location` empty when you need broad geographic coverage.
- Use a broader location when a narrow geographic term produces limited results.
- Add `excludeWords` carefully because a matching exclusion term causes the complete result description to be skipped.
- Review the output before using collected contact information in important workflows.
- Treat email addresses as publicly discovered data that may require independent validation.

### Troubleshooting

**Empty Results**

Try broader or more specific keywords, remove an overly narrow location, or add additional email-domain suffixes. A search may also have limited publicly indexed information.

**Fewer Emails Than `maxEmails`**

The limit is a target rather than a guarantee. The selected searches may simply contain fewer matching publicly indexed email addresses.

**Unexpectedly Missing Results**

Check whether the result description contains one of the configured exclusion words or phrases. Also verify that the email uses one of the configured domain suffixes.

**Location Produces Limited Coverage**

Try a broader country, state, or city value, or leave `location` empty when geographic filtering is not required.

**Invalid Input**

Check that `keywords`, `customDomains`, and `excludeWords` are arrays of strings and that `maxEmails` is an integer between 1 and 10,000.

**Partial Results**

Partial collection can occur when suitable indexed results are limited or when a search does not provide enough matching information. Review the query configuration before increasing collection targets.

### Frequently Asked Questions

**What does the SimilarWeb Email Scraper do?**

It searches for publicly indexed SimilarWeb results using your configured keywords and email-domain suffixes, then extracts matching email addresses from result descriptions.

**What data does the SimilarWeb Email Scraper return?**

Each record contains `keyword`, `title`, `description`, `url`, and `email`.

**Can I use multiple keywords?**

Yes. The `keywords` input is an array, so you can provide multiple search terms in a single run.

**Can I target a specific country or city?**

Yes. Enter a country, state, or city in the optional `location` field to narrow the search geographically.

**Can I search for different email providers?**

Yes. The `customDomains` field accepts multiple email-domain suffixes, such as `@gmail.com`, `@yahoo.com`, and `@outlook.com`.

**How does `maxEmails` work?**

The configured value is applied to each keyword + domain combination. It can be set from 1 through 10,000 in the Actor input schema. Free users have an effective ceiling of 100 when a higher or unspecified value is supplied.

**Can I exclude unwanted results?**

Yes. Add words or phrases to `excludeWords`. If a configured exclusion matches the result description, that description is skipped and no email is collected from it.

**Why did the Actor return fewer emails than requested?**

`maxEmails` controls the collection target; it does not guarantee that the requested number exists in publicly indexed results. Broader keywords, additional domains, or a broader location can increase potential coverage.

**Is the output structured?**

Yes. Results are stored as structured Apify dataset records with consistent fields for keyword, title, description, URL, and email.

**Can duplicate email addresses appear?**

The Actor keeps track of collected addresses and does not add the same email address repeatedly during the run.

### NLP Keywords

- SimilarWeb email scraper
- SimilarWeb email extraction
- SimilarWeb contact data
- SimilarWeb email finder
- SimilarWeb contact scraper
- SimilarWeb business contacts
- SimilarWeb lead data
- SimilarWeb profile emails
- SimilarWeb email collection
- SimilarWeb data extraction
- SimilarWeb contact discovery
- SimilarWeb business research
- SimilarWeb lead research
- SimilarWeb email addresses
- SimilarWeb structured data
- SimilarWeb search results
- SimilarWeb company research
- SimilarWeb keyword search
- SimilarWeb contact information
- SimilarWeb dataset

### Related Keywords

- SimilarWeb email scraper tool
- scrape emails from SimilarWeb
- SimilarWeb email extractor
- SimilarWeb contact information scraper
- SimilarWeb business email finder
- SimilarWeb lead scraper
- SimilarWeb public email scraper
- SimilarWeb company email extraction
- SimilarWeb email data collection
- SimilarWeb contact discovery tool
- SimilarWeb keyword email search
- SimilarWeb business contact extractor
- SimilarWeb profile email extraction
- SimilarWeb email research
- SimilarWeb lead generation research
- SimilarWeb contact dataset
- SimilarWeb email search automation
- SimilarWeb company contact data
- SimilarWeb business research scraper
- SimilarWeb email address finder

### Final Overview

The **SimilarWeb Email Scraper** provides a configurable way to discover publicly indexed email addresses associated with SimilarWeb search results. Users can supply multiple keywords, optionally narrow searches by location, select email-domain suffixes, configure collection targets, and filter descriptions using exclusion terms.

The resulting Apify dataset preserves the keyword, result title, description, URL, and matching email address, giving users useful context for business research, lead research, company research, market analysis, and structured contact-data workflows.

For better coverage, start with focused but varied keywords, use appropriate email domains, and adjust location filtering according to the research objective. Always review and validate important contact information before relying on it for downstream activities.

*Contact me:* <Alphascraper69@gmail.com>

# Actor input Schema

## `keywords` (type: `array`):

A list of keywords or queries to search for.

## `location` (type: `string`):

Optional country, state or city used to narrow the search. Leave it empty to search without a geographic filter.

## `customDomains` (type: `array`):

List of custom email domains

## `maxEmails` (type: `integer`):

How many addresses each search keyword + domain suffix combination may collect before the finder moves on to the next one. This is a per-combination target, not a run-wide total: with 3 keyword and 2 Domains and a limit of 20, the run works through all 6 combinations and aims for up to 20 addresses in each, so up to 120 overall. Lower values finish sooner and cost less; higher values dig deeper but never guarantee a fuller result, since the run can only find what is publicly listed.

## `excludeWords` (type: `array`):

Words or phrases you do not want to see.

## Actor input object example

```json
{
  "keywords": [
    "marketing manager",
    "online business"
  ],
  "location": "",
  "customDomains": [
    "@gmail.com"
  ],
  "maxEmails": 5,
  "excludeWords": []
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "marketing manager",
        "online business"
    ],
    "location": "",
    "customDomains": [
        "@gmail.com"
    ],
    "excludeWords": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("email_scraper/similarweb-email-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": [
        "marketing manager",
        "online business",
    ],
    "location": "",
    "customDomains": ["@gmail.com"],
    "excludeWords": [],
}

# Run the Actor and wait for it to finish
run = client.actor("email_scraper/similarweb-email-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "marketing manager",
    "online business"
  ],
  "location": "",
  "customDomains": [
    "@gmail.com"
  ],
  "excludeWords": []
}' |
apify call email_scraper/similarweb-email-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,email_scraper/similarweb-email-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ChVgh1t24M4cqWhJz/builds/ubNf88hjSw21Ppicg/openapi.json
