# Hugging Face Email Scraper (`leads-scraper/hugging-face-email-scraper`) Actor

Hugging Face Email Scraper SD - Hugging Face Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from Hugging Face results by keyword, location and email domain - Hugging Face email extractor.

- **URL**: https://apify.com/leads-scraper/hugging-face-email-scraper.md
- **Developed by:** [Leads Scraper](https://apify.com/leads-scraper) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.49 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### Hugging Face Email Scraper - extract model author emails from Hugging Face results

The **Hugging Face Email Scraper** is an Apify Actor that collects publicly indexed contact emails tied to Hugging Face accounts, model cards, dataset cards and Spaces. You supply keywords, optional location and the email domains you care about, and it returns a clean dataset of ML leads.

Machine-learning practitioners publish addresses in places the Hugging Face Email Scraper is built to reach: the "Citation" and "Contact" sections of a model card, the maintainer block of a dataset card, the README of a Space, and organisation pages for research labs.

This is the tool people mean when they search for a Hugging Face model author email, a way to contact an ML researcher about a checkpoint, or a Hugging Face email extractor that produces a spreadsheet instead of one address at a time.

#### What the Hugging Face Email Scraper actually reads

The Hugging Face Email Scraper reads **only what Google has already indexed** on `huggingface.co`. It builds Google queries with the `site:` operator, fetches result pages through the Apify GOOGLE_SERP proxy, and extracts emails from titles and snippets.

It does **not** log into the Hub, call the Hugging Face API, use `huggingface_hub`, or download a repository. There is no browser, no JavaScript rendering, no authentication and no cookies.

That distinction matters here. A contact address inside a model card is reachable only because Google indexed that card - not because the Actor cloned the repo or read the model's files.

#### Who the Hugging Face Email Scraper is built for

ML platform vendors selling to model authors, research labs recruiting engineers, compute and inference providers building a launch list, and academic teams looking for collaborators all need the same thing: a repeatable way to turn a research area into contactable authors.

The Hugging Face Email Scraper turns "speech recognition, Gmail addresses" into rows you can filter, dedupe and export as CSV, JSON or Excel, or pull straight from the Apify API.

***

### Key features of the Hugging Face Email Scraper

| Feature | What it does |
|---|---|
| Google `site:` dorking | Queries `huggingface.co` through the Apify GOOGLE_SERP proxy |
| Query expansion | Runs base, quoted and `intitle:` variants plus one variant per query modifier; base queries run first |
| Domain-filtered extraction | Keeps only emails ending in your `customDomains` values |
| Global deduplication | One row per unique email across every query and every page of the run |
| Obfuscation handling | Understands `name [at] domain [dot] com`, `name (at) domain`, `name @ domain.com`, `domain .com`, zero-width characters and the full-width `＠` |
| Junk filter | Rejects placeholders such as `email@`, `yourname@`, `test@`, `xxx@` and single-character locals |
| Boundary-correct matching | `@gmail.com` never matches inside `@gmail.company` or `@gmail.com.br` |
| Soft-wrap repair | Drops a hit that is only the tail of another email in the same result block |
| Structural parsing | Locates the `<h3>` title then the smallest surrounding block, instead of relying on Google's CSS class names |
| Handle resolution | Recovers the `huggingface.co/{username}` handle from user, organisation and repo URLs |
| Whole-page fallback | If Google's markup changes, the run degrades to "emails without account details" rather than "no emails" |
| Block detection | CAPTCHA, "unusual traffic" and consent pages are detected and retried, not counted as empty |
| Retries and backoff | Up to 3 attempts per page with exponential backoff and a fresh proxy session per request |
| Requeue of failures | Blocked or failed queries are re-queued once at the end of the run |
| Async concurrency | An `asyncio` worker pool with a shared stop signal on `maxEmails` |
| Resumable state | Progress is stored in the key-value store keyed by a hash of the input, with throttled saves plus saves on `PERSIST_STATE`, `MIGRATING` and `ABORTING` |
| Streaming output | Every lead is pushed to the dataset immediately, so partial results survive an abort |
| Run summary | Logs pages fetched, blocked pages, retries and emails per page |

***

### How the Hugging Face Email Scraper works

The pipeline behind the Hugging Face Email Scraper is deliberately simple, which is why it runs quickly and is hard to break.

1. The Hugging Face Email Scraper reads your input: keywords, location, email domains and limits.
2. It builds Google queries with the `site:` operator, for example `site:huggingface.co model author "@gmail.com"`.
3. It fetches Google result pages asynchronously through the Apify GOOGLE_SERP proxy using `aiohttp`.
4. It parses each result block structurally, finding the `<h3>` title and the smallest block around it.
5. It extracts emails from the block text with a domain-filtered regex.
6. It deduplicates globally and pushes each lead straight to the Apify dataset.

#### Why query expansion matters for ML researcher outreach

Google caps a single query at roughly 300 results, so one query has a hard ceiling no matter how many pages you allow.

Query expansion is how the Hugging Face Email Scraper gets past it: each keyword is combined with each email domain in several phrasings, and every variant carries its own result budget.

The default `queryModifiers` are `email`, `contact`, `maintainer`, `author` and `support` - the words that actually appear beside an address in a model card or a dataset card.

#### What the Hugging Face Email Scraper does not do

It does not log into the Hub, use the Hugging Face API or `huggingface_hub`, clone repositories, read gated-model access requests, or reach anything that is not already public in Google's index.

It is an independent Apify Actor and is not affiliated with, supported by or endorsed by Hugging Face.

***

### Hugging Face Email Scraper input fields

Every field the Hugging Face Email Scraper accepts is optional except `keywords`. Defaults are exactly those shipped in the Actor's input schema.

| Field | Type | Default | Meaning |
|---|---|---|---|
| `keywords` | array (required) | `["machine learning", "model author"]` | Search terms describing the Hugging Face accounts you want (research area, task, role) |
| `location` | string | `""` | Optional location phrase added to every query |
| `customDomains` | array | `["@gmail.com", "@yahoo.com"]` | Only emails on these domains are kept; the leading `@` is optional |
| `maxEmails` | integer 1-10000 | `20` | Stop after this many unique emails |
| `countryCode` | string | `""` | Two-letter country for the search proxy (US, GB, DE...) |
| `expandQueries` | boolean | `true` | Search each keyword x domain pair in several phrasings |
| `queryModifiers` | array | `["email", "contact", "maintainer", "author", "support"]` | Extra words combined with each keyword when expansion is on |
| `maxPagesPerQuery` | integer 1-50 | `30` | Page cap per query |
| `maxConcurrency` | integer 1-20 | `5` | Parallel queries |

#### Example Hugging Face Email Scraper input

```json
{
  "keywords": ["speech recognition model", "vision language model", "dataset maintainer"],
  "location": "",
  "customDomains": ["@gmail.com", "@ed.ac.uk"],
  "maxEmails": 500,
  "countryCode": "GB",
  "expandQueries": true,
  "queryModifiers": ["email", "contact", "maintainer", "author", "support"],
  "maxPagesPerQuery": 30,
  "maxConcurrency": 5
}
```

#### Tuning the Hugging Face Email Scraper for better yield

Task and architecture language works best in the Hugging Face Email Scraper. `speech recognition`, `retrieval augmented generation`, `quantized llm`, `instruction tuned` and `image segmentation dataset` match how model and dataset cards are written far better than `researcher`.

Academic domains are a strong lever here. Adding university domains to `customDomains` alongside `@gmail.com` reaches lab authors that consumer-domain filters miss entirely.

Leave `expandQueries` on. Turning it off gives one query per keyword and domain, which is faster but caps out quickly.

***

### Hugging Face Email Scraper output fields

Every dataset item the Hugging Face Email Scraper produces carries all 14 fields below. Fields are never dropped; they are empty or `null` when Google did not expose the value.

| Field | Meaning |
|---|---|
| `network` | Platform name |
| `keyword` | The keyword that produced the lead |
| `query` | The exact Google query used |
| `title` | Raw result title |
| `accountName` | Account label Google prints (user handle, organisation, or display name) |
| `fullName` | Display name parsed from a profile-style title; empty for model or dataset card titles |
| `username` | URL-safe Hub handle when one is exposed; otherwise `null` |
| `profileUrl` | Canonical `https://huggingface.co/{username}` URL when a handle is known; otherwise empty |
| `url` | Direct platform link when exposed, else the profile URL |
| `description` | Bio or card snippet, cleaned of labels and engagement counters |
| `email` | Lower-cased email address |
| `emailDomain` | The matched domain, for example `@gmail.com` |
| `possiblyTruncated` | `true` when Google's snippet ellipsis touched the email - verify before sending |
| `foundAt` | ISO 8601 UTC timestamp |

#### Example Hugging Face Email Scraper output

```json
[
  {
    "network": "Hugging Face",
    "keyword": "speech recognition model",
    "query": "site:huggingface.co speech recognition model contact \"@gmail.com\"",
    "title": "ana-vidal/whisper-pt-small · Hugging Face",
    "accountName": "ana-vidal",
    "fullName": "",
    "username": "ana-vidal",
    "profileUrl": "https://huggingface.co/ana-vidal",
    "url": "https://huggingface.co/ana-vidal/whisper-pt-small",
    "description": "Fine-tuned Whisper for European Portuguese. Questions and collaboration: ana.vidal.asr@gmail.com",
    "email": "ana.vidal.asr@gmail.com",
    "emailDomain": "@gmail.com",
    "possiblyTruncated": false,
    "foundAt": "2026-08-31T09:14:22.481Z"
  },
  {
    "network": "Hugging Face",
    "keyword": "dataset maintainer",
    "query": "site:huggingface.co intitle:\"dataset\" maintainer \"@ed.ac.uk\"",
    "title": "northlab/clinical-notes-benchmark · Datasets at Hugging Face",
    "accountName": "northlab",
    "fullName": "",
    "username": "northlab",
    "profileUrl": "https://huggingface.co/northlab",
    "url": "https://huggingface.co/datasets/northlab/clinical-notes-benchmark",
    "description": "Benchmark for clinical note summarisation. Maintainer: j.okafor@ed.ac.uk. Please cite the accompanying ...",
    "email": "j.okafor@ed.ac.uk",
    "emailDomain": "@ed.ac.uk",
    "possiblyTruncated": true,
    "foundAt": "2026-08-31T09:14:29.113Z"
  },
  {
    "network": "Hugging Face",
    "keyword": "vision language model",
    "query": "site:huggingface.co vision language model author email \"@gmail.com\"",
    "title": "Marek Kowalski (marekk) - Hugging Face",
    "accountName": "marekk",
    "fullName": "Marek Kowalski",
    "username": "marekk",
    "profileUrl": "https://huggingface.co/marekk",
    "url": "https://huggingface.co/marekk",
    "description": "Multimodal research. VLM checkpoints and eval harnesses. Reach me at marek.kowalski.vlm@gmail.com",
    "email": "marek.kowalski.vlm@gmail.com",
    "emailDomain": "@gmail.com",
    "possiblyTruncated": false,
    "foundAt": "2026-08-31T09:15:03.902Z"
  }
]
```

Note how the Hugging Face Email Scraper fills `url` and `profileUrl` differently on the first two rows: `url` points at the model or dataset repository, while `profileUrl` resolves to the author or organisation account.

***

### Use cases for the Hugging Face Email Scraper

| Use case | How the Hugging Face Email Scraper helps |
|---|---|
| ML talent sourcing | Build a list of practitioners by task, modality and framework, with a Hub profile to review first |
| Research collaboration | Reach the authors of checkpoints and benchmarks adjacent to your own work |
| ML tooling and infrastructure sales | Contact people already shipping models when launching inference, eval or annotation products |
| Dataset licensing and provenance | Write to dataset maintainers about terms, corrections or reuse |
| Academic and lab outreach | Filter to university domains and reach research groups publishing on the Hub |
| Developer relations | Assemble a launch list of model authors for a beta, grant or credits programme |
| Ecosystem research | Measure how many authors in a task area publish a public contact channel at all |
| Conference and workshop programmes | Shortlist speakers and reviewers from a research keyword |

#### Reaching model authors well

Model authors get a lot of noise, so use Hugging Face Email Scraper output carefully. Read the `description` snippet and open `url` before writing: naming the actual checkpoint, its base model or its eval results changes the tone of the whole message.

Remember that many addresses on model cards exist for citation and reproducibility questions. A pitch that ignores that context reads as spam even when the address is genuinely public.

#### Covering the wider ML footprint

Researchers rarely publish in one place. Pair the Hugging Face Email Scraper with the [GitHub Email Scraper](https://apify.com/leads-scraper/github-email-scraper) for training code and with the **PyPI Email Scraper** for the Python packages that wrap it.

***

### Expected results from the Hugging Face Email Scraper

In live test runs against Google, roughly **8 out of 10 parsed Hugging Face results carried an account identity** - a Hub handle, an organisation name, or both.

That is one of the stronger ratios in this family, because Hub URLs are structured as `huggingface.co/{owner}/{repo}` and Google prints the owner in the title. Handle resolution therefore works well on model and dataset cards, not only on profiles.

Hugging Face Email Scraper rows without a handle still carry `email`, `emailDomain`, `url`, `description` and `query`. Volume always depends on your keywords, domains and location - no yield is guaranteed.

***

### Limitations of the Hugging Face Email Scraper

Read this before planning a campaign around Hugging Face Email Scraper output.

- **Publicly indexed emails only.** The Hugging Face Email Scraper finds an address only when it is already visible in Google's index. Addresses behind gated-model access forms, private repos or unindexed cards are out of reach.
- **No Hub API and no repository access.** Model files, config JSON and `huggingface_hub` responses are never read. Only Google result titles and snippets are parsed.
- **Google's ~300-result cap.** A single query returns roughly 300 results maximum, which is exactly why `expandQueries` exists and why turning it off reduces volume.
- **`possiblyTruncated`.** When this is `true`, Google's snippet ellipsis touched the email and it may be cut off. Verify those rows before sending anything.
- **`username` and `profileUrl` can be empty.** They are populated only when the Hub exposes a handle in the result; some Spaces and blog-style pages show only a display name, leaving `accountName` filled but `username` `null`.
- **Apify GOOGLE_SERP proxy required.** The Actor cannot run without Apify proxy credentials.
- **Free plan cap.** Free Apify plans are limited to 100 emails per run. Paid plans are uncapped.
- **Variable results.** Output depends on keywords, domains, location and Google's index, which changes over time. Two runs a month apart will not be identical.

***

### Responsible use

Addresses on model and dataset cards are published for research correspondence - citation questions, reproducibility issues, licensing and collaboration. Treat them that way.

If you contact people in the EU or UK, GDPR applies: have a lawful basis, identify yourself, say where you found the address, and honour opt-outs immediately. Respect Hugging Face community norms and any stated contact preference on a card or profile.

Academic addresses deserve extra care. Many are institutional, monitored, and covered by policies that make bulk commercial mail a genuine problem for the recipient.

***

### Hugging Face Email Scraper FAQ

#### Does the Hugging Face Email Scraper use the Hugging Face API?

No. It parses Google search results for pages on `huggingface.co` that Google has already indexed. There is no API call, no `huggingface_hub` usage, no login and no repository download.

#### Does it need a Hugging Face account or token?

No. There is no authentication, no browser, no JavaScript rendering and no cookies. The only credential required is your Apify account's GOOGLE_SERP proxy access.

#### Where do the emails the Hugging Face Email Scraper returns come from?

From the text Google prints in result titles and snippets - model card contact and citation sections, dataset card maintainer blocks, Space READMEs, and profile or organisation bios.

#### Can the Hugging Face Email Scraper find authors of gated models?

Only if an address for them is already public in Google's index. It cannot submit access requests or read anything behind a gate.

#### How many emails can the Hugging Face Email Scraper return?

`maxEmails` accepts 1 to 10000 and defaults to 20. Free Apify plans are capped at 100 emails per run; paid plans are uncapped. Volume depends on your keywords and domains.

#### Why is `username` sometimes `null`?

Because Google did not expose a Hub handle for that result. The row still carries `email`, `accountName`, `url` and `description`. On this platform it is the minority case, since owner handles appear in most Hub URLs.

#### What does `possiblyTruncated: true` mean?

Google's snippet ellipsis touched the email, so the address may be incomplete. Verify those rows before you use them.

#### Can I target university domains with the Hugging Face Email Scraper?

Yes, and it works well here. Put them in `customDomains`, for example `["@ed.ac.uk", "@mit.edu"]`. Only emails ending in those domains are kept, and the leading `@` is optional.

#### Does it handle obfuscated addresses like `name [at] gmail [dot] com`?

Yes. The extractor normalises `[at]`, `(at)`, spaced `@`, spaced `.com`, zero-width characters and the full-width `＠`, and it filters placeholders such as `test@` and `yourname@`.

#### Can I run it for a specific country?

Set `countryCode` to a two-letter code such as `US`, `GB` or `DE` to steer the search proxy, and put a city, region or institution in `location` to append it to every query.

#### What happens if a Hugging Face Email Scraper run is interrupted?

Progress is saved in the key-value store keyed by a hash of your input, with saves on `PERSIST_STATE`, `MIGRATING` and `ABORTING`. Leads are pushed to the dataset as they are found, so nothing already collected is lost.

#### Is the Hugging Face Email Scraper affiliated with Hugging Face?

No. It is an independent Apify Actor, and it is not supported, endorsed or affiliated with Hugging Face.

***

### Related Actors

The Hugging Face Email Scraper belongs to a Developer & Technology family of Apify Actors that all work the same way on different platforms. Run several to cover an ecosystem instead of a single site.

| Actor | What it collects |
|---|---|
| [Hugging Face Email and Phone Number Scraper](https://apify.com/neuro-scraper/hugging-face-email-and-phone-number-scraper) | Emails and phone numbers from Hugging Face |
| [Hugging Face Phone Number Scraper](https://apify.com/neuro-scraper/hugging-face-phone-number-scraper) | Public phone numbers from Hugging Face |
| [App Store Email Scraper](https://apify.com/leads-scraper/app-store-email-scraper) | Public contact emails from App Store |
| [Atlassian Marketplace Email Scraper](https://apify.com/leads-scraper/atlassian-marketplace-email-scraper) | Public contact emails from Atlassian Marketplace |
| [Bitbucket Email Scraper](https://apify.com/neuro-scraper/bitbucket-email-scraper) | Public contact emails from Bitbucket |
| [Chrome Web Store Email Scraper](https://apify.com/leads-scraper/chrome-web-store-email-scraper) | Public contact emails from Chrome Web Store |
| [CodePen Email Scraper](https://apify.com/neuro-scraper/codepen-email-scraper) | Public contact emails from CodePen |
| [Confluence Email Scraper](https://apify.com/neuro-scraper/confluence-email-scraper) | Public contact emails from Confluence |
| [Dev.to Email Scraper](https://apify.com/neuro-scraper/dev-to-email-scraper) | Public contact emails from DEV Community |
| [Docker Hub Email Scraper](https://apify.com/neuro-scraper/docker-hub-email-scraper) | Public contact emails from Docker Hub |
| [Figma Community Email Scraper](https://apify.com/neuro-scraper/figma-community-email-scraper) | Public contact emails from Figma Community |
| [Firefox Add-ons Email Scraper](https://apify.com/neuro-scraper/firefox-add-ons-email-scraper) | Public contact emails from Firefox Add-ons |
| [GitHub Email Scraper](https://apify.com/leads-scraper/github-email-scraper) | Public contact emails from GitHub |
| [GitLab Email Scraper](https://apify.com/neuro-scraper/gitlab-email-scraper) | Public contact emails from GitLab |
| [Google Play Email Scraper](https://apify.com/leads-scraper/google-play-email-scraper) | Public contact emails from Google Play |
| [Hashnode Email Scraper](https://apify.com/neuro-scraper/hashnode-email-scraper) | Public contact emails from Hashnode |
| [HubSpot Marketplace Email Scraper](https://apify.com/leads-scraper/hubspot-marketplace-email-scraper) | Public contact emails from HubSpot Marketplace |
| [Jira Email Scraper](https://apify.com/neuro-scraper/jira-email-scraper) | Public contact emails from Jira |
| [Maven Central Email Scraper](https://apify.com/neuro-scraper/maven-central-email-scraper) | Public contact emails from Maven Central |
| [Microsoft AppSource Email Scraper](https://apify.com/neuro-scraper/microsoft-appsource-email-scraper) | Public contact emails from Microsoft AppSource |
| [Salesforce AppExchange Email Scraper](https://apify.com/leads-scraper/salesforce-appexchange-email-scraper) | Public contact emails from Salesforce AppExchange |
| [Shopify App Store Email Scraper](https://apify.com/leads-scraper/shopify-app-store-email-scraper) | Public contact emails from Shopify App Store |
| [Slack App Directory Email Scraper](https://apify.com/leads-scraper/slack-app-directory-email-scraper) | Public contact emails from Slack App Directory |
| [SourceForge Email Scraper](https://apify.com/leads-scraper/sourceforge-email-scraper) | Public contact emails from SourceForge |
| [Stack Overflow Email Scraper](https://apify.com/leads-scraper/stack-overflow-email-scraper) | Public contact emails from Stack Overflow |
| [Unity Asset Store Email Scraper](https://apify.com/leads-scraper/unity-asset-store-email-scraper) | Public contact emails from Unity Asset Store |
| [Unreal Engine Marketplace Email Scraper](https://apify.com/leads-scraper/unreal-engine-marketplace-email-scraper) | Public contact emails from Unreal Engine Marketplace |
| [WordPress Plugin Directory Email Scraper](https://apify.com/neuro-scraper/wordpress-plugin-directory-email-scraper) | Public contact emails from WordPress Plugin Directory |
| [WordPress Theme Directory Email Scraper](https://apify.com/leads-scraper/wordpress-theme-directory-email-scraper) | Public contact emails from WordPress Theme Directory |
| [Zapier App Directory Email Scraper](https://apify.com/leads-scraper/zapier-app-directory-email-scraper) | Public contact emails from Zapier App Directory |
| [App Store Email and Phone Number Scraper](https://apify.com/leads-scraper/app-store-email-and-phone-number-scraper) | Emails and phone numbers from App Store |
| [Atlassian Marketplace Email and Phone Number Scraper](https://apify.com/neuro-scraper/atlassian-marketplace-email-and-phone-number-scraper) | Emails and phone numbers from Atlassian Marketplace |
| [Bitbucket Email and Phone Number Scraper](https://apify.com/neuro-scraper/bitbucket-email-and-phone-number-scraper) | Emails and phone numbers from Bitbucket |
| [Chrome Web Store Email and Phone Number Scraper](https://apify.com/neuro-scraper/chrome-web-store-email-and-phone-number-scraper) | Emails and phone numbers from Chrome Web Store |
| [CodePen Email and Phone Number Scraper](https://apify.com/neuro-scraper/codepen-email-and-phone-number-scraper) | Emails and phone numbers from CodePen |
| [Confluence Email and Phone Number Scraper](https://apify.com/neuro-scraper/confluence-email-and-phone-number-scraper) | Emails and phone numbers from Confluence |
| [DEV Community Email and Phone Number Scraper](https://apify.com/neuro-scraper/dev-to-email-and-phone-number-scraper) | Emails and phone numbers from DEV Community |
| [Docker Hub Email and Phone Number Scraper](https://apify.com/neuro-scraper/docker-hub-email-and-phone-number-scraper) | Emails and phone numbers from Docker Hub |
| [Figma Community Email and Phone Number Scraper](https://apify.com/neuro-scraper/figma-community-email-and-phone-number-scraper) | Emails and phone numbers from Figma Community |
| [Firefox Add-ons Email and Phone Number Scraper](https://apify.com/neuro-scraper/firefox-add-ons-email-and-phone-number-scraper) | Emails and phone numbers from Firefox Add-ons |
| [GitHub Email and Phone Number Scraper](https://apify.com/leads-scraper/github-email-and-phone-number-scraper) | Emails and phone numbers from GitHub |
| [GitLab Email and Phone Number Scraper](https://apify.com/neuro-scraper/gitlab-email-and-phone-number-scraper) | Emails and phone numbers from GitLab |

### Leave a review

If the Hugging Face Email Scraper saved you time, please leave a star rating and a short review
on the Actor page.

Reviews are how other buyers judge whether a tool works, and they tell us which features to
build next.

If something did not work, email <neurodata.apify@gmail.com>
instead - bugs get fixed faster than they get complained about.

### Support

Questions, a bug, or a custom build of the Hugging Face Email Scraper for your own domain list? Email **neurodata.apify@gmail.com** and we will get back to you.

# Actor input Schema

## `keywords` (type: `array`):

Search terms describing the Hugging Face accounts you want (niche, job title, industry).

## `location` (type: `string`):

Optional location phrase added to every query (e.g. "New York").

## `customDomains` (type: `array`):

Only emails ending with one of these domains are collected. With or without the leading @. Each domain is searched separately, so more domains means more results but a longer run - remove some for a faster, narrower search, or add your own (e.g. @company.com).

## `maxEmails` (type: `integer`):

Stop once this many unique emails have been collected.

## `countryCode` (type: `string`):

Two-letter country code for the search proxy (e.g. US, GB, DE). Empty for any.

## `expandQueries` (type: `boolean`):

Search each keyword x domain pair with several phrasings. Recommended - Google caps a single query at ~300 results.

## `queryModifiers` (type: `array`):

Extra words combined with each keyword when Expand queries is on. Tuned for Hugging Face.

## `maxPagesPerQuery` (type: `integer`):

Google rarely returns more than ~30 pages for one query.

## `maxConcurrency` (type: `integer`):

How many queries run in parallel.

## Actor input object example

```json
{
  "keywords": [
    "machine learning",
    "model author"
  ],
  "location": "",
  "customDomains": [
    "@gmail.com",
    "@yahoo.com"
  ],
  "maxEmails": 20,
  "countryCode": "",
  "expandQueries": true,
  "queryModifiers": [
    "email",
    "contact",
    "maintainer",
    "author",
    "support"
  ],
  "maxPagesPerQuery": 30,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

Records produced by Hugging Face Email Scraper, stored in the run's default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "machine learning",
        "model author"
    ],
    "location": "",
    "customDomains": [
        "@gmail.com",
        "@yahoo.com"
    ],
    "countryCode": "",
    "queryModifiers": [
        "email",
        "contact",
        "maintainer",
        "author",
        "support"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("leads-scraper/hugging-face-email-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": [
        "machine learning",
        "model author",
    ],
    "location": "",
    "customDomains": [
        "@gmail.com",
        "@yahoo.com",
    ],
    "countryCode": "",
    "queryModifiers": [
        "email",
        "contact",
        "maintainer",
        "author",
        "support",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("leads-scraper/hugging-face-email-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "machine learning",
    "model author"
  ],
  "location": "",
  "customDomains": [
    "@gmail.com",
    "@yahoo.com"
  ],
  "countryCode": "",
  "queryModifiers": [
    "email",
    "contact",
    "maintainer",
    "author",
    "support"
  ]
}' |
apify call leads-scraper/hugging-face-email-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,leads-scraper/hugging-face-email-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/S7Rv4lmNQYUekxZ5I/builds/vKtxC1flyp5UCT3Av/openapi.json
