Kaggle Email Scraper
Pricing
from $2.49 / 1,000 results
Kaggle Email Scraper
Kaggle Email Scraper SD - Kaggle Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from Kaggle results by keyword, location and email domain - Kaggle email extractor.
Pricing
from $2.49 / 1,000 results
Rating
0.0
(0)
Developer
Neuro Scraper
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
7 days ago
Last modified
Categories
Share
Kaggle Email Scraper - extract public data scientist and ML practitioner emails from Kaggle
The Kaggle Email Scraper collects publicly indexed contact emails from Kaggle: data scientists, machine learning engineers, notebook authors, dataset publishers and competition participants who left an address in public text.
You supply keywords, an optional location and a list of email domains. The Kaggle Email Scraper builds domain-scoped Google queries against kaggle.com, parses each result block and writes every matching address into a structured dataset.
There is no login, no Kaggle API, no headless browser and no cookies. This Kaggle email extractor reads only what Google has already indexed publicly.
One honest caveat up front. Kaggle yields less cleanly than creator platforms, because a large share of indexed kaggle.com pages are datasets, competitions and notebooks rather than user profiles, so a fair proportion of rows arrive with an email but no username and no profile URL.
That is worth knowing before you plan a run. The addresses are still real and usable; you simply get fewer rows where a handle and a kaggle.com profile URL come attached.
Technical recruiters, ML tooling vendors, developer-relations teams, bootcamps and research groups use the Kaggle Email Scraper for contact discovery across the data science community.
Features of the Kaggle Email Scraper
Everything below is implemented in the Kaggle Email Scraper today. No roadmap items, no aspirational claims.
| Feature | What it means in practice |
|---|---|
Google site: search automation | Every query is scoped to kaggle.com, so results stay on-platform |
| Query expansion | Base, quoted and intitle: variants plus one variant per query modifier |
| Recruiting-tuned modifiers | Defaults are email, contact, inquiries, hire, work with me |
| Domain-filtered extraction | Only emails on your customDomains list are kept |
| Global deduplication | Each unique address reaches the dataset exactly once across all queries and pages |
| Obfuscation-aware parser | Understands name [at] domain [dot] com, name (at) domain, name @ domain.com, domain .com |
| Unicode normalisation | Handles zero-width characters and the full-width at sign |
| Junk filter | Rejects placeholders such as email@, yourname@, test@, xxx@ and single-character locals |
| Boundary-correct matching | @gmail.com will not match inside @gmail.company or @gmail.com.br |
| Soft-wrap repair | Discards a hit that is only the tail of another email in the same result block |
| Concurrency control | An asyncio worker pool runs queries in parallel with a shared stop signal on maxEmails |
| Retry logic | Up to 3 attempts per page with exponential backoff and a fresh proxy session per request |
| Block detection | CAPTCHA, "unusual traffic" and consent pages are detected and retried, never scored as empty |
| Re-queue pass | Blocked or failed queries are retried once more at the end of the run |
| Resumable state | Progress persists in the key-value store, keyed by a hash of your input |
| Structural parsing | Result blocks are located via the <h3> title, not by Google's CSS class names |
| Fallback parser | A Google markup change degrades to "emails without account details", not "no emails" |
| Streaming dataset writes | Leads appear in the dataset immediately, so exports can begin mid-run |
How the Kaggle Email Scraper works
The Kaggle Email Scraper is a search-results crawler. It never opens kaggle.com, never downloads a dataset and never executes a notebook.
1. Read input. Keywords, location, email domains, page caps and concurrency are validated first.
2. Build queries. Queries use the site: operator, for example:
site:kaggle.com machine learning engineer "@gmail.com" "Bangalore".
3. Expand queries. With expandQueries on, each keyword and domain pair is issued as a base query, a quoted variant, an intitle: variant and one variant per entry in queryModifiers. Base queries always run first.
4. Fetch pages. The Kaggle Email Scraper fetches result pages asynchronously with aiohttp through the Apify GOOGLE_SERP proxy, paginating up to maxPagesPerQuery.
5. Parse result blocks. The parser locates the <h3> title, then the smallest surrounding block, so titles, snippets and site labels stay attached to the correct record.
6. Extract and normalise. A domain-filtered regex pulls candidates from the block text, normalises them, and removes junk, tails and boundary mismatches.
7. Deduplicate and store. Every new address is deduplicated globally and pushed straight into the Apify dataset as structured data.
The Kaggle Email Scraper run summary logs pages fetched, blocked pages, retries and emails per page, which is especially useful here for comparing profile-shaped keywords against notebook-shaped ones.
What data does the Kaggle Email Scraper extract?
Each row is one email paired with whatever metadata Google exposed alongside it.
You get the address, the matched domain, the account or page label, a parsed name where the title looks like a profile, a handle and profile URL where Kaggle exposes one, a cleaned snippet and full provenance.
Provenance matters. The keyword and query fields tell you which phrasing surfaced this record, which is how you steer the next Kaggle Email Scraper run toward profile pages rather than dataset pages.
When a handle is available the Kaggle Email Scraper rebuilds profileUrl as https://www.kaggle.com/<handle>. When Google returns a dataset, competition or notebook page instead, username will be null and profileUrl empty.
Input schema for the Kaggle Email Scraper
Every Kaggle Email Scraper input field is listed below with the exact name, type and default from the Actor input schema.
| Field | Type | Default | Description |
|---|---|---|---|
keywords | array (required) | ["data scientist", "machine learning"] | Search terms describing the Kaggle accounts you want (niche, job title, industry) |
location | string | "" | Optional location phrase added to every query, e.g. "New York" |
customDomains | array | ["@gmail.com", "@yahoo.com"] | Only emails ending with one of these domains are collected; leading @ optional |
maxEmails | integer 1-10000 | 20 | Stop once this many unique emails have been collected |
countryCode | string | "" | Two-letter country code for the search proxy (US, GB, DE). Empty for any |
expandQueries | boolean | true | Search each keyword and domain pair with several phrasings |
queryModifiers | array | ["email", "contact", "inquiries", "hire", "work with me"] | Extra words combined with each keyword when expansion is on |
maxPagesPerQuery | integer 1-50 | 30 | Page cap per query |
maxConcurrency | integer 1-20 | 5 | How many queries run in parallel |
Input example
{"keywords": ["data scientist", "machine learning engineer", "computer vision"],"location": "","customDomains": ["@gmail.com", "@outlook.com"],"maxEmails": 300,"countryCode": "US","expandQueries": true,"queryModifiers": ["email", "contact", "inquiries", "hire", "work with me"],"maxPagesPerQuery": 30,"maxConcurrency": 5}
Output schema of the Kaggle Email Scraper
Every dataset item the Kaggle Email Scraper produces carries all fourteen fields below.
| Field | Meaning |
|---|---|
network | Platform name |
keyword | The keyword that produced the lead |
query | The exact Google query used |
title | Raw result title |
accountName | Account label Google prints (handle, display name or page label) |
fullName | Display name parsed from a profile-style title; empty for dataset and notebook titles |
username | URL-safe handle when Kaggle exposes one; otherwise null |
profileUrl | Canonical account URL when a handle is known; otherwise empty |
url | Direct Kaggle link when exposed, else the profile URL |
description | Bio or page snippet, cleaned of labels and engagement counters |
email | Lower-cased email address |
emailDomain | The matched domain, e.g. @gmail.com |
possiblyTruncated | true when Google's snippet ellipsis touched the email - verify before sending |
foundAt | ISO 8601 UTC timestamp |
Output example
{"network": "Kaggle","keyword": "computer vision","query": "site:kaggle.com computer vision \"@gmail.com\" contact","title": "Ravi Deshmukh | Notebooks Expert | Kaggle","accountName": "ravideshmukh","fullName": "Ravi Deshmukh","username": "ravideshmukh","profileUrl": "https://www.kaggle.com/ravideshmukh","url": "https://www.kaggle.com/ravideshmukh","description": "Computer vision and medical imaging. Open to consulting and research collaborations. Contact: ravi.deshmukh.ml@gmail.com","email": "ravi.deshmukh.ml@gmail.com","emailDomain": "@gmail.com","possiblyTruncated": false,"foundAt": "2026-08-31T14:12:09.348Z"}
The Kaggle Email Scraper dataset exports as JSON, CSV, Excel or XML from the Apify console, or streams through the Apify API into your ATS, CRM or outreach tool.
How to use the Kaggle Email Scraper
Step 1 - bias your keywords toward people, not projects. Terms like data scientist, machine learning engineer, NLP researcher or Kaggle Master bring back more profile pages; terms like titanic dataset bring back competition pages with no owner attached.
Step 2 - pick sensible email domains. @gmail.com dominates among individual practitioners. Add academic or corporate domains if you are targeting a specific institution.
Step 3 - start small. Run the Kaggle Email Scraper with maxEmails at 20 to 50 first, then inspect how many rows came back with a populated username.
Step 4 - keep the recruiting-tuned modifiers. hire, contact, inquiries and work with me match the phrasing people use when they are open to work.
Step 5 - add location for geography-bound hiring. Practitioners sometimes state a city or country, but many do not, so expect location filtering to reduce volume noticeably.
Step 6 - clean before you send. Filter out rows where possiblyTruncated is true, split rows with and without a username, then run the remaining Kaggle Email Scraper results through email verification.
Use cases for the Kaggle Email Scraper
| Use case | Who runs it | Typical keywords |
|---|---|---|
| Data science and ML recruiting | Technical recruiters and in-house sourcers | data scientist, machine learning engineer, kaggle master |
| Developer tools and MLOps prospecting | Founders and sales teams at ML vendors | deep learning, pytorch, feature engineering |
| Developer relations and community building | DevRel and community managers | notebook, tutorial, open source |
| Research collaboration outreach | Labs, universities and research groups | nlp researcher, computer vision, time series |
| Bootcamp and course enrolment | Education providers and training platforms | beginner, learning data science, portfolio |
| Freelance and contract data work | Agencies and consultancies | freelance data scientist, analytics consultant |
| Conference and hackathon recruiting | Event organisers and sponsors | competition, hackathon, benchmark |
| Talent-market research | Analysts studying the ML labour market | any role keyword, varied by location |
Kaggle is a demonstrated-skill community rather than a resume site, which is why sourcers value it, and why the Kaggle Email Scraper is best used as one input among several rather than a standalone list.
For academic contacts the ResearchGate Email Scraper and the Academia.edu Email Scraper are usually richer, and practitioners who publish courses often appear in the Coursera Email Scraper or the Pluralsight Email Scraper.
Why choose this Kaggle Email Scraper
Manual sourcing on Kaggle means paging through leaderboards, opening profiles, reading bios and copying addresses. The Kaggle Email Scraper automates that loop and returns a deduplicated dataset.
Every row is auditable. The query and keyword fields show exactly how a lead was found, which matters more here than elsewhere because it lets you separate profile-shaped queries from dataset-shaped ones.
The extraction layer is deliberately conservative. Junk filtering, boundary-correct domain matching and soft-wrap repair mean far fewer garbage rows to clean downstream.
Long Kaggle Email Scraper runs survive interruption. State persists in the key-value store and saves fire on Apify's PERSIST_STATE, MIGRATING and ABORTING events.
Worked example: a computer vision hiring pipeline
You are hiring three computer vision engineers and want a top-of-funnel list of 300 practitioners.
Set keywords to ["computer vision", "image segmentation", "kaggle master"], customDomains to ["@gmail.com"], countryCode to "US" and maxEmails to 300, leaving expandQueries on.
The Kaggle Email Scraper issues base, quoted and intitle: variants plus one query per modifier, paginating each independently and streaming leads as they are found.
Export the Kaggle Email Scraper dataset to CSV, drop possiblyTruncated rows, then sort by whether username is populated so your sourcers work the identifiable profiles first.
Limitations you should know before running
The Kaggle Email Scraper is deliberately transparent about its boundaries. Read these before planning a large data collection run.
Weaker identity resolution than creator platforms. Many indexed kaggle.com pages are datasets, competitions and notebooks rather than user profiles, so a fair share of Kaggle Email Scraper rows carry an email with no handle attached.
Only publicly indexed emails. The Kaggle Email Scraper returns addresses that already appear in Google's index. Practitioners who never published an address will not surface.
Google's per-query ceiling. A single query caps out at roughly 300 results. Query expansion exists to work around that ceiling, which is why it defaults to on.
possiblyTruncated. When Google's snippet ellipsis touches an address, this field is true and the email may be incomplete. Verify those rows before sending.
Apify GOOGLE_SERP proxy required. The Kaggle Email Scraper cannot run without Apify proxy credentials that include the GOOGLE_SERP group.
Free-plan cap. Running the Kaggle Email Scraper on a free Apify plan limits results to 100 emails per run. Paid plans are uncapped.
Handle availability. username and profileUrl are populated only when Google's result exposes a Kaggle handle. Rows without one still carry accountName and, sometimes, fullName. This is a Google limitation, not a bug.
No guaranteed volume. Kaggle Email Scraper yield varies with keywords, domains and location. No result count is promised.
The Kaggle Email Scraper is not affiliated with, endorsed by or officially supported by Kaggle.
Related Actors
Every Actor below runs the same engine as the Kaggle Email Scraper, with identical input fields and an identical output schema; only the target platform differs.
| Actor | What it collects |
|---|---|
| Kaggle Email and Phone Number Scraper | Emails and phone numbers from Kaggle |
| Kaggle Phone Number Scraper | Public phone numbers from Kaggle |
| Academia.edu Email Scraper | Public contact emails from Academia.edu |
| Buy Me a Coffee Email Scraper | Public contact emails from Buy Me a Coffee |
| Carrd Email Scraper | Public contact emails from Carrd |
| Domestika Email Scraper | Public contact emails from Domestika |
| edX Email Scraper | Public contact emails from edX |
| Gumroad Email Scraper | Public contact emails from Gumroad |
| Kajabi Email Scraper | Public contact emails from Kajabi |
| Ko-fi Email Scraper | Public contact emails from Ko-fi |
| Linktree Email Scraper | Public contact emails from Linktree |
| MasterClass Email Scraper | Public contact emails from MasterClass |
| Meetup Email Scraper | Public contact emails from Meetup |
| Pluralsight Email Scraper | Public contact emails from Pluralsight |
| Preply Email Scraper | Public contact emails from Preply |
| ResearchGate Email Scraper | Public contact emails from ResearchGate |
| Skillshare Email Scraper | Public contact emails from Skillshare |
| Teachable Email Scraper | Public contact emails from Teachable |
| Udemy Email Scraper | Public contact emails from Udemy |
| Academia.edu Email and Phone Number Scraper | Emails and phone numbers from Academia.edu |
| Buy Me a Coffee Email and Phone Number Scraper | Emails and phone numbers from Buy Me a Coffee |
| Carrd Email and Phone Number Scraper | Emails and phone numbers from Carrd |
| Domestika Email and Phone Number Scraper | Emails and phone numbers from Domestika |
| edX Email and Phone Number Scraper | Emails and phone numbers from edX |
| Gumroad Email and Phone Number Scraper | Emails and phone numbers from Gumroad |
| Kajabi Email and Phone Number Scraper | Emails and phone numbers from Kajabi |
| Ko-fi Email and Phone Number Scraper | Emails and phone numbers from Ko-fi |
| Linktree Email and Phone Number Scraper | Emails and phone numbers from Linktree |
| MasterClass Email and Phone Number Scraper | Emails and phone numbers from MasterClass |
| Meetup Email and Phone Number Scraper | Emails and phone numbers from Meetup |
| Pluralsight Email and Phone Number Scraper | Emails and phone numbers from Pluralsight |
| Preply Email and Phone Number Scraper | Emails and phone numbers from Preply |
| ResearchGate Email and Phone Number Scraper | Emails and phone numbers from ResearchGate |
| Skillshare Email and Phone Number Scraper | Emails and phone numbers from Skillshare |
| Teachable Email and Phone Number Scraper | Emails and phone numbers from Teachable |
| Academia.edu Phone Number Scraper | Public phone numbers from Academia.edu |
| Buy Me a Coffee Phone Number Scraper | Public phone numbers from Buy Me a Coffee |
| Carrd Phone Number Scraper | Public phone numbers from Carrd |
| Domestika Phone Number Scraper | Public phone numbers from Domestika |
| edX Phone Number Scraper | Public phone numbers from edX |
| Gumroad Phone Number Scraper | Public phone numbers from Gumroad |
| Kajabi Phone Number Scraper | Public phone numbers from Kajabi |
FAQ about the Kaggle Email Scraper
What is the Kaggle Email Scraper?
It is an Apify Actor that extracts publicly indexed contact emails from Kaggle by running domain-scoped Google searches and parsing the resulting search-result blocks.
Does the Kaggle Email Scraper log into Kaggle or use its API?
No. It never logs in, never calls a Kaggle API and never opens a browser, so there are no cookies and no authentication anywhere in the pipeline.
Why do so many rows have no username?
Because much of Kaggle's indexed surface is datasets, competitions and notebooks rather than profiles. Those pages have no handle for the Kaggle Email Scraper to resolve, so username is null and profileUrl is empty.
Can I improve the share of rows with a handle?
Often, yes. Role-shaped keywords such as data scientist, machine learning engineer or kaggle master return more profile pages than dataset-shaped keywords.
Where do the emails actually come from?
From publicly indexed titles, snippets and site labels, typically bios and descriptions where someone published an address themselves.
What does possiblyTruncated: true mean?
Google's snippet ellipsis touched the address, so it may be cut off. Verify those rows before contacting anyone.
How many emails will one run of the Kaggle Email Scraper return?
maxEmails accepts 1 to 10000, but free Apify plans are capped at 100 emails per run and paid plans are uncapped. Real volume depends on keywords, domains and location.
Is a proxy required?
Yes. The Kaggle Email Scraper requires the Apify GOOGLE_SERP proxy and cannot run without Apify proxy credentials.
Should I leave expandQueries on?
Almost always. Google returns roughly 300 results per query, and expansion multiplies your reachable surface through base, quoted, intitle: and modifier variants.
Can I target a specific country?
Yes. Set countryCode for the search proxy and optionally add a location phrase, keeping in mind that many practitioners never state a city.
What happens if my run is interrupted or migrated?
State persists in the key-value store keyed by a hash of your input, with throttled saves plus saves on PERSIST_STATE, MIGRATING and ABORTING, so an interrupted run resumes instead of restarting.
Can I email everyone in the dataset immediately?
Treat the Kaggle Email Scraper output as raw lead data. Deduplication is automatic, but verification, consent and compliance with GDPR, CAN-SPAM and Kaggle's terms remain your responsibility.
Leave a review
If the Kaggle Email Scraper saved you time, please leave a star rating and a short review on the Actor page.
Reviews are how other buyers judge whether a tool works, and they tell us which features to build next.
If something did not work, email neurodata.apify@gmail.com instead - bugs get fixed faster than they get complained about.
Support
Questions about the Kaggle Email Scraper, bug reports or a custom build? Email neurodata.apify@gmail.com.