# Google Maps Email Extractor — Emails, Phones & Socials ($1/1K) (`datamech/google-maps-email-extractor`) Actor

Google Maps email extractor and lead generator: get emails, phone numbers, socials (Facebook, Instagram, LinkedIn, WhatsApp) and contact pages from business websites, plus a website audit (dead/parked, HTTPS, SSL, mobile, CMS). $1 per 1,000 sites.

- **URL**: https://apify.com/datamech/google-maps-email-extractor.md
- **Developed by:** [Fer](https://apify.com/datamech) (community)
- **Categories:** Lead generation, Marketing, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 business website analyzeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google Maps Email Extractor: Leads, Contact Details, Phones, Socials & Website Audit

**Google Maps Email Extractor** (Google Maps Leads) turns a list of business websites, or the results of our Google Maps scraper, into a ready-to-use lead list. For every business it visits the website and returns **emails, phone numbers, social media profiles (Facebook, Instagram, LinkedIn, X/Twitter, TikTok, YouTube, WhatsApp) and the contact page**, plus a **website health audit**: is the site dead or parked, does it use HTTPS, is the SSL certificate valid, is it mobile-friendly, which CMS it runs on (WordPress, Wix, Shopify, Squarespace...) and how fast it responds.

Use it as a **Google Maps email extractor**, a **contact details scraper** (emails, phone numbers, socials) or a **website audit** tool. It is built for **lead generation**, **cold email outreach** and **web agencies that sell websites**: find local businesses on Google Maps, get their contact details, and instantly see which ones have a broken, insecure or outdated website.

Works out of the box for **English, Spanish and Portuguese** sites (US, Mexico, Latin America, Spain, Brazil), and also understands German and French contact pages (Kontakt, Impressum, Contact, À propos).

### What can this Google Maps email extractor do?

- 📧 **Find emails** from the homepage and the most likely contact pages (Contact, About, Impressum, Contacto, Nosotros, Quiénes somos, Contato, Fale conosco, Kontakt...). Decodes Cloudflare-protected emails (`[email protected]`), `mailto:` links, schema.org/JSON-LD data and `info (at) domain (dot) com` style obfuscation.
- 🧹 **Clean, deduplicated emails**: syntax-validated, junk removed (example.com, tracking/error-reporting addresses, image file names like `logo@2x.png`, placeholder addresses), and the business's own domain ranked first.
- 📞 **Phone numbers** from `tel:` links, structured data and the page text, validated and formatted internationally (`+52 33 1527 7234`, `+1 402 709 7059`).
- 🔗 **Social profiles**: Facebook, Instagram, LinkedIn, X/Twitter, TikTok, YouTube and **WhatsApp** (`wa.me` links, very common for Latin American businesses). Share buttons and website-builder footer links are ignored.
- 🩺 **Website audit for agencies**: reachable or dead, parked/for-sale domains, HTTP status, redirects, HTTPS and SSL errors, mobile viewport, page title and meta description, detected platform/CMS, response time.
- 🗺️ **Google Maps integration**: enrich a dataset from our Google Maps scraper, paste places directly, or let this Actor run the Google Maps search for you. All place fields (name, placeId, address, phone, categories, rating...) are kept next to the enrichment.
- 💸 **Cheap and light**: plain HTTP requests (no browser), runs in 256 MB of memory, a handful of requests per website.

### How to use it

#### 1. Enrich a list of websites

Paste the websites into **Business websites** (plain domains like `acme-plumbing.com` work too) and click **Start**.

```json
{
    "startUrls": [
        { "url": "https://inlawplumbing.com/" },
        { "url": "https://yucatandental.com/" }
    ],
    "maxPagesPerSite": 5
}
```

#### 2. Enrich results of our Google Maps scraper

Run our [Google Maps Scraper](https://apify.com/datamech/apify-google-maps-scraper), copy the **dataset ID** of the run and paste it into **Google Maps dataset ID**. Every place becomes one row, with all its Maps fields kept.

```json
{
    "datasetId": "YOUR_GOOGLE_MAPS_DATASET_ID",
    "onlyWithEmail": true
}
```

You can also pass places directly in the `places` field (only `website` is required), which is handy from the API or from integrations.

#### 3. Search Google Maps and enrich in one run

Fill in **Search Google Maps first** and **Search location**. The Actor starts our Google Maps scraper, waits for it, and enriches every place it finds.

```json
{
    "searchStringsArray": ["dentista", "plomero"],
    "locationQuery": "Guadalajara, Jalisco",
    "maxPlaces": 50,
    "mapsLanguage": "es",
    "includePlacesWithoutWebsite": true
}
```

The Google Maps run is a separate run on your account and is billed by the Google Maps scraper.

### Input options

| Option | Default | What it does |
| --- | --- | --- |
| `startUrls` / `websites` | 3 example sites | Websites to enrich. |
| `datasetId` / `places` | – | Google Maps places to enrich. |
| `searchStringsArray`, `locationQuery`, `maxPlaces`, `mapsLanguage` | – / – / 20 / `en` | Run the Google Maps search first. |
| `includePlacesWithoutWebsite` | `false` | Also output places without a website (`siteStatus: no_website`), great for agencies. |
| `maxPagesPerSite` | 5 | Homepage + up to 4 contact-like pages. Stops early once contact page, email and phone are found. |
| `maxConcurrency` | 30 | Websites crawled in parallel. Automatically lowered when memory is tight, so the default is safe at 256 MB. |
| `requestTimeoutSecs` | 15 | Timeout per page. |
| `siteTimeoutSecs` | 60 | Maximum time spent on one website (all its pages). |
| `maxRetries` | 1 | Retries for the homepage on network errors, 5xx, 429 and 403. |
| `onlyWithEmail` | `false` | Save only businesses with at least one email. |
| `skipDeadSites` | `false` | Do not save dead, parked or missing websites. |
| `includeBlockedSites` | `false` | Save captcha/WAF blocked sites too. By default they are listed in `FAILED_SITES` and **not charged**. |
| `proxyConfiguration` | none | Optional. First pass uses this proxy if you set it (off by default). Blocked sites are retried via SHADER → Apify automatic (no groups, FREE datacenter) → none. Residential is never selected automatically. Your input, if set, is used as-is. |

### Output example

One row per business. Example from a real run:

```json
{
    "name": "Yucatan Dental",
    "website": "https://yucatandental.com/",
    "domain": "yucatandental.com",
    "primaryEmail": "info@yucatandental.com",
    "emails": ["info@yucatandental.com", "yucatandental@gmail.com"],
    "phones": ["+52 999 924 9895"],
    "facebook": "https://www.facebook.com/yucatandental",
    "instagram": null,
    "linkedin": null,
    "twitter": null,
    "tiktok": null,
    "youtube": null,
    "whatsapp": null,
    "contactPageUrl": "https://yucatandental.com/contact-us/",
    "siteStatus": "ok",
    "reachable": true,
    "httpStatus": 200,
    "finalUrl": "https://yucatandental.com/",
    "redirected": false,
    "redirectedToOtherDomain": false,
    "https": true,
    "sslValid": true,
    "sslError": null,
    "responseTimeMs": 241,
    "hasMobileViewport": true,
    "hasTitle": true,
    "pageTitle": "Inicio - Yucatan Dental",
    "hasMetaDescription": true,
    "metaDescription": "Yucatan Dental is a small professional dental practice located in the Centro Historico of Merida, Mexico...",
    "platform": "WordPress (WooCommerce)",
    "language": "en",
    "parkedReason": null,
    "hasContactForm": true,
    "pagesCrawled": 2,
    "error": null,
    "checkedAt": "2026-10-08T04:20:29.990Z"
}
```

When the input comes from Google Maps, the place fields (`title`, `placeId`, `address`, `city`, `phone`, `categories`, `rating`, `reviewsCount`, `googleUrl`...) are added to the same row. Heavy fields such as reviews and opening hours are left out.

#### Website status values

| `siteStatus` | Meaning |
| --- | --- |
| `ok` | Website loads normally. |
| `dead` | Domain does not resolve, refuses connections or times out. |
| `parked` | Parked or for-sale domain, suspended hosting account, registrar landing page or bare server page ("Index of /", "It works!"). See `parkedReason`. |
| `blocked` | Site is online but refuses automated visits (403/429 or a bot challenge). |
| `http_error` | Homepage returns an error such as 404 or 500. |
| `social_only` | The listed website is a social profile or a WhatsApp chat link (`wa.me`, `wa.link`, `api.whatsapp.com`, `chat.whatsapp.com`). The link (and the phone, when the URL contains it) is saved; WhatsApp's own page is not crawled. |
| `directory_listing` | The listed website is a profile on a directory, booking or delivery platform (Doctoralia, Yelp, Tripadvisor, Linktree, OpenTable, Booksy, Uber Eats, Rappi, Zocdoc, Healthgrades...). The platform's own phones, emails and social accounts are not returned. |
| `no_website` | Google Maps place without any website (only with `includePlacesWithoutWebsite`). |

You can download results as **CSV, Excel, JSON or XML**, or use the **Leads overview**, **Contacts** and **Website audit** views in the Output tab.

### Pricing

This Actor uses **pay-per-event** pricing: **$1.00 per 1,000 websites analyzed** ($0.001 per business row saved to the dataset), plus a tiny $0.00005 fee per run start.

- You only pay for rows saved to the dataset. Filtered rows (`onlyWithEmail`, `skipDeadSites`), invalid URLs, and captcha/WAF blocked sites (unless you turn on `includeBlockedSites`) are **free**. Dead and parked sites are saved, because they are useful leads for web agencies.
- Example: enriching 1,000 Google Maps places costs about **$1.00**. No proxy, browser or compute costs are added on top.
- The run stops automatically, with a clear status message, when it reaches your **maximum number of results** or your **maximum cost per run**.
- Each run saves a `RUN_SUMMARY` record in the key-value store with counts, email/phone/social hit rates, skipped rows and the stop reason.
- When you use **Search Google Maps first**, the Google Maps scraper run is billed separately by that Actor.

### Performance

Designed to run at **256 MB**. Pages are fetched over HTTP (no browser), bodies keep the first ~1.35 MB plus the last 150 KB (so footer mailto links are not cut), and parallelism backs off when memory is tight. Shared-host IPs are limited to 2 concurrent requests; SiteGround (detected from `host-header` / `sg-captcha` / CNAME / `35.214.0.0/16` and `35.215.0.0/16`) is limited to **1** in-flight request so a large run does not trip their volume captcha. Bot-challenge pages (SiteGround 202, Cloudflare, Sucuri) are **not charged** by default: they go to the `FAILED_SITES` key-value record and are retried at the end of the run through Apify Proxy (SHADER, then automatic/no-groups for FREE plans, new session per site). A further 45 s concurrency-1 retry runs only when a proxy is actually available. `RUN_SUMMARY.retryProxy` records which proxy was used and `blockedRecovered` how many sites came back. `primaryEmail` is the best guess: same-domain addresses first, then free-provider addresses; legal, privacy, hiring and abuse mailboxes are kept but ranked last. Internal, localhost and private-IP URLs are rejected. Each run writes `RUN_SUMMARY` even if the run is aborted or is about to time out.

On 1,000 mixed websites at 256 MB (concurrency 30): about 160 s (~6 sites/s), 159 MB peak RSS, 0 OOM. On 50 real dentist and plumber sites: about 6 s, 56% with email, 78% with phone, 74% with a social profile.

### Spanish and Latin America support

The crawler recognizes Spanish and Portuguese contact pages (*Contacto, Contáctanos, Nosotros, Quiénes somos, Ubicación, Sucursales, Contato, Fale conosco, Quem somos*), Spanish obfuscation (`hola (arroba) dominio (punto) mx`), WhatsApp links, and formats Mexican and other Latin American phone numbers correctly even when the site has no country hints — including pre-2019 Mexican mobiles written as `+52 1 …`, `044`/`045` or `01`. It also detects platforms common in the region such as Tiendanube, VTEX and Site123.

### FAQ

**Is it legal to extract emails from websites?**
The Actor only reads publicly available business websites. Contact details of businesses are generally public, but personal data may be protected by laws such as GDPR or CAN-SPAM. Use the data responsibly and check the rules in your country before sending marketing emails.

**Why didn't it find an email for some businesses?**
Many small businesses publish only a contact form or a phone number. The Actor checks the homepage and the most likely contact pages, but it does not fill in forms, run JavaScript, or guess addresses. In our test on 50 US and Mexican dentists and plumbers, about 56% of businesses had an email on their website.

**Does it work with websites built with Wix, Squarespace or other JavaScript-heavy builders?**
Usually yes: most builders still include contact details in the HTML. Pages that load everything with JavaScript after the page opens may return fewer results.

**Some sites show as blocked. Can I still get them?**
Those sites returned a bot challenge (SiteGround captcha, Cloudflare 403, …). By default they are **not saved and not charged**; they are listed in the `FAILED_SITES` key-value record and retried through Apify Proxy. Turn on `includeBlockedSites` if you want the empty rows in the dataset anyway.

**Can I use it without Google Maps?**
Yes. Any list of websites works: from a CRM, a directory export or a spreadsheet.

**How do I connect it to Google Sheets, Make, Zapier or my CRM?**
Use the Apify integrations or the API. Example API call:

```bash
curl -X POST "https://api.apify.com/v2/acts/datamech~google-maps-email-extractor/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"websites": ["inlawplumbing.com", "yucatandental.com"], "onlyWithEmail": true}'
```

**Do I need a proxy?**
No for the first pass on normal lists. Blocked SiteGround/Cloudflare sites are retried automatically through Apify Proxy when the account has datacenter access (SHADER, or automatic mode on the FREE plan). You can also set `proxyConfiguration` yourself; that value is respected. Residential proxy is never selected unless you choose it.

### Feedback

Found a website where an email, phone or social profile was missed or wrong? Open an issue in the Issues tab with the URL, and we will improve the extractor.

# Actor input Schema

## `startUrls` (type: `array`):

Websites to enrich, one per business (homepage URLs work best). Plain domains like <code>acme-plumbing.com</code> are accepted. Leave empty if you use Google Maps places below.

## `websites` (type: `array`):

Same as <b>Business websites</b> but as a simple list of strings. Handy when calling the Actor through the API.

## `datasetId` (type: `string`):

ID of a dataset produced by our Google Maps scraper (or any dataset whose items have a <code>website</code> field). Each place becomes one output row.

## `places` (type: `array`):

Place objects pasted directly, e.g. <code>\[{"title": "In-Law Plumbing", "website": "https://inlawplumbing.com", "placeId": "..."}]</code>. Only <code>website</code> is required.

## `searchStringsArray` (type: `array`):

Optional. If set, the Actor first runs our Google Maps scraper with these search terms (e.g. <code>dentist</code>, <code>plomero</code>), then enriches every place found. The Maps run is billed separately on your account by the Google Maps scraper.

## `locationQuery` (type: `string`):

Location appended to each search term, e.g. <code>Austin, TX</code> or <code>Guadalajara, Jalisco</code>.

## `maxPlaces` (type: `integer`):

Maximum number of Google Maps places per search term.

## `mapsLanguage` (type: `string`):

Language code for the Google Maps search, e.g. <code>en</code>, <code>es</code>, <code>pt-BR</code>.

## `includePlacesWithoutWebsite` (type: `boolean`):

Output places that have no website (marked <code>siteStatus: no_website</code>). Useful for agencies selling websites. Each row counts as a result.

## `maxPagesPerSite` (type: `integer`):

Homepage plus the most likely contact pages (Contact, About, Impressum, Contacto, Nosotros, Contato, Kontakt, À propos...). The crawl of a site stops early once a contact page, an email and a phone were found.

## `maxConcurrency` (type: `integer`):

How many websites are crawled in parallel. Pages of one website are always fetched one after another to be polite. The Actor automatically lowers parallelism when memory gets tight, so the default is safe at 256 MB.

## `requestTimeoutSecs` (type: `integer`):

Maximum time for downloading one page. Slower sites are reported as unreachable.

## `siteTimeoutSecs` (type: `integer`):

Maximum total time spent on one website (all its pages, retries and http/https fallbacks). When it runs out, the row is saved with whatever was found so far.

## `maxRetries` (type: `integer`):

Retries for the homepage on timeouts, resets and 5xx/429 responses. Contact pages are not retried.

## `onlyWithEmail` (type: `boolean`):

Save only businesses where at least one email was found. Filtered rows are not saved and not charged.

## `skipDeadSites` (type: `boolean`):

Do not save businesses whose website is unreachable, parked/for sale, or missing. Filtered rows are not saved and not charged. Captcha/WAF blocked sites are omitted by default even when this is off (see <code>includeBlockedSites</code>).

## `includeBlockedSites` (type: `boolean`):

By default, sites that only returned a bot challenge (SiteGround 202, Cloudflare 403, etc.) are not saved and not charged. They are listed in the <code>FAILED_SITES</code> key-value record. Turn this on to save those rows anyway.

## `proxyConfiguration` (type: `object`):

Optional. First-pass crawling uses this proxy if you set it (off by default, to keep normal sites cheap). Whatever you pick here is respected, including residential. If you leave it off, blocked sites are retried with a new session through: (1) Apify datacenter group SHADER, (2) Apify Proxy automatic / no groups (FREE-plan datacenter), (3) no proxy. Residential is never selected automatically. Without a working proxy the slow 45 s retry is skipped.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://inlawplumbing.com/"
    },
    {
      "url": "https://www.ashrowandental.com/"
    },
    {
      "url": "https://yucatandental.com/"
    }
  ],
  "maxPlaces": 20,
  "mapsLanguage": "en",
  "includePlacesWithoutWebsite": false,
  "maxPagesPerSite": 5,
  "maxConcurrency": 30,
  "requestTimeoutSecs": 15,
  "siteTimeoutSecs": 60,
  "maxRetries": 1,
  "onlyWithEmail": false,
  "skipDeadSites": false,
  "includeBlockedSites": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `leads` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

## `failedSites` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://inlawplumbing.com/"
        },
        {
            "url": "https://www.ashrowandental.com/"
        },
        {
            "url": "https://yucatandental.com/"
        }
    ],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("datamech/google-maps-email-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [
        { "url": "https://inlawplumbing.com/" },
        { "url": "https://www.ashrowandental.com/" },
        { "url": "https://yucatandental.com/" },
    ],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("datamech/google-maps-email-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://inlawplumbing.com/"
    },
    {
      "url": "https://www.ashrowandental.com/"
    },
    {
      "url": "https://yucatandental.com/"
    }
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call datamech/google-maps-email-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datamech/google-maps-email-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/kAa9GKu3PsbHYfb2K/builds/WBXdEK2YMXH7t13sW/openapi.json
