Kaggle Email Scraper avatar

Kaggle Email Scraper

Pricing

from $2.49 / 1,000 results

Go to Apify Store
Kaggle Email Scraper

Kaggle Email Scraper

Kaggle Email Scraper SD - Kaggle Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from Kaggle results by keyword, location and email domain - Kaggle email extractor.

Pricing

from $2.49 / 1,000 results

Rating

0.0

(0)

Developer

Neuro Scraper

Neuro Scraper

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

7 days ago

Last modified

Categories

Share

Kaggle Email Scraper - extract public data scientist and ML practitioner emails from Kaggle

The Kaggle Email Scraper collects publicly indexed contact emails from Kaggle: data scientists, machine learning engineers, notebook authors, dataset publishers and competition participants who left an address in public text.

You supply keywords, an optional location and a list of email domains. The Kaggle Email Scraper builds domain-scoped Google queries against kaggle.com, parses each result block and writes every matching address into a structured dataset.

There is no login, no Kaggle API, no headless browser and no cookies. This Kaggle email extractor reads only what Google has already indexed publicly.

One honest caveat up front. Kaggle yields less cleanly than creator platforms, because a large share of indexed kaggle.com pages are datasets, competitions and notebooks rather than user profiles, so a fair proportion of rows arrive with an email but no username and no profile URL.

That is worth knowing before you plan a run. The addresses are still real and usable; you simply get fewer rows where a handle and a kaggle.com profile URL come attached.

Technical recruiters, ML tooling vendors, developer-relations teams, bootcamps and research groups use the Kaggle Email Scraper for contact discovery across the data science community.


Features of the Kaggle Email Scraper

Everything below is implemented in the Kaggle Email Scraper today. No roadmap items, no aspirational claims.

FeatureWhat it means in practice
Google site: search automationEvery query is scoped to kaggle.com, so results stay on-platform
Query expansionBase, quoted and intitle: variants plus one variant per query modifier
Recruiting-tuned modifiersDefaults are email, contact, inquiries, hire, work with me
Domain-filtered extractionOnly emails on your customDomains list are kept
Global deduplicationEach unique address reaches the dataset exactly once across all queries and pages
Obfuscation-aware parserUnderstands name [at] domain [dot] com, name (at) domain, name @ domain.com, domain .com
Unicode normalisationHandles zero-width characters and the full-width at sign
Junk filterRejects placeholders such as email@, yourname@, test@, xxx@ and single-character locals
Boundary-correct matching@gmail.com will not match inside @gmail.company or @gmail.com.br
Soft-wrap repairDiscards a hit that is only the tail of another email in the same result block
Concurrency controlAn asyncio worker pool runs queries in parallel with a shared stop signal on maxEmails
Retry logicUp to 3 attempts per page with exponential backoff and a fresh proxy session per request
Block detectionCAPTCHA, "unusual traffic" and consent pages are detected and retried, never scored as empty
Re-queue passBlocked or failed queries are retried once more at the end of the run
Resumable stateProgress persists in the key-value store, keyed by a hash of your input
Structural parsingResult blocks are located via the <h3> title, not by Google's CSS class names
Fallback parserA Google markup change degrades to "emails without account details", not "no emails"
Streaming dataset writesLeads appear in the dataset immediately, so exports can begin mid-run

How the Kaggle Email Scraper works

The Kaggle Email Scraper is a search-results crawler. It never opens kaggle.com, never downloads a dataset and never executes a notebook.

1. Read input. Keywords, location, email domains, page caps and concurrency are validated first.

2. Build queries. Queries use the site: operator, for example: site:kaggle.com machine learning engineer "@gmail.com" "Bangalore".

3. Expand queries. With expandQueries on, each keyword and domain pair is issued as a base query, a quoted variant, an intitle: variant and one variant per entry in queryModifiers. Base queries always run first.

4. Fetch pages. The Kaggle Email Scraper fetches result pages asynchronously with aiohttp through the Apify GOOGLE_SERP proxy, paginating up to maxPagesPerQuery.

5. Parse result blocks. The parser locates the <h3> title, then the smallest surrounding block, so titles, snippets and site labels stay attached to the correct record.

6. Extract and normalise. A domain-filtered regex pulls candidates from the block text, normalises them, and removes junk, tails and boundary mismatches.

7. Deduplicate and store. Every new address is deduplicated globally and pushed straight into the Apify dataset as structured data.

The Kaggle Email Scraper run summary logs pages fetched, blocked pages, retries and emails per page, which is especially useful here for comparing profile-shaped keywords against notebook-shaped ones.


What data does the Kaggle Email Scraper extract?

Each row is one email paired with whatever metadata Google exposed alongside it.

You get the address, the matched domain, the account or page label, a parsed name where the title looks like a profile, a handle and profile URL where Kaggle exposes one, a cleaned snippet and full provenance.

Provenance matters. The keyword and query fields tell you which phrasing surfaced this record, which is how you steer the next Kaggle Email Scraper run toward profile pages rather than dataset pages.

When a handle is available the Kaggle Email Scraper rebuilds profileUrl as https://www.kaggle.com/<handle>. When Google returns a dataset, competition or notebook page instead, username will be null and profileUrl empty.


Input schema for the Kaggle Email Scraper

Every Kaggle Email Scraper input field is listed below with the exact name, type and default from the Actor input schema.

FieldTypeDefaultDescription
keywordsarray (required)["data scientist", "machine learning"]Search terms describing the Kaggle accounts you want (niche, job title, industry)
locationstring""Optional location phrase added to every query, e.g. "New York"
customDomainsarray["@gmail.com", "@yahoo.com"]Only emails ending with one of these domains are collected; leading @ optional
maxEmailsinteger 1-1000020Stop once this many unique emails have been collected
countryCodestring""Two-letter country code for the search proxy (US, GB, DE). Empty for any
expandQueriesbooleantrueSearch each keyword and domain pair with several phrasings
queryModifiersarray["email", "contact", "inquiries", "hire", "work with me"]Extra words combined with each keyword when expansion is on
maxPagesPerQueryinteger 1-5030Page cap per query
maxConcurrencyinteger 1-205How many queries run in parallel

Input example

{
"keywords": ["data scientist", "machine learning engineer", "computer vision"],
"location": "",
"customDomains": ["@gmail.com", "@outlook.com"],
"maxEmails": 300,
"countryCode": "US",
"expandQueries": true,
"queryModifiers": ["email", "contact", "inquiries", "hire", "work with me"],
"maxPagesPerQuery": 30,
"maxConcurrency": 5
}

Output schema of the Kaggle Email Scraper

Every dataset item the Kaggle Email Scraper produces carries all fourteen fields below.

FieldMeaning
networkPlatform name
keywordThe keyword that produced the lead
queryThe exact Google query used
titleRaw result title
accountNameAccount label Google prints (handle, display name or page label)
fullNameDisplay name parsed from a profile-style title; empty for dataset and notebook titles
usernameURL-safe handle when Kaggle exposes one; otherwise null
profileUrlCanonical account URL when a handle is known; otherwise empty
urlDirect Kaggle link when exposed, else the profile URL
descriptionBio or page snippet, cleaned of labels and engagement counters
emailLower-cased email address
emailDomainThe matched domain, e.g. @gmail.com
possiblyTruncatedtrue when Google's snippet ellipsis touched the email - verify before sending
foundAtISO 8601 UTC timestamp

Output example

{
"network": "Kaggle",
"keyword": "computer vision",
"query": "site:kaggle.com computer vision \"@gmail.com\" contact",
"title": "Ravi Deshmukh | Notebooks Expert | Kaggle",
"accountName": "ravideshmukh",
"fullName": "Ravi Deshmukh",
"username": "ravideshmukh",
"profileUrl": "https://www.kaggle.com/ravideshmukh",
"url": "https://www.kaggle.com/ravideshmukh",
"description": "Computer vision and medical imaging. Open to consulting and research collaborations. Contact: ravi.deshmukh.ml@gmail.com",
"email": "ravi.deshmukh.ml@gmail.com",
"emailDomain": "@gmail.com",
"possiblyTruncated": false,
"foundAt": "2026-08-31T14:12:09.348Z"
}

The Kaggle Email Scraper dataset exports as JSON, CSV, Excel or XML from the Apify console, or streams through the Apify API into your ATS, CRM or outreach tool.


How to use the Kaggle Email Scraper

Step 1 - bias your keywords toward people, not projects. Terms like data scientist, machine learning engineer, NLP researcher or Kaggle Master bring back more profile pages; terms like titanic dataset bring back competition pages with no owner attached.

Step 2 - pick sensible email domains. @gmail.com dominates among individual practitioners. Add academic or corporate domains if you are targeting a specific institution.

Step 3 - start small. Run the Kaggle Email Scraper with maxEmails at 20 to 50 first, then inspect how many rows came back with a populated username.

Step 4 - keep the recruiting-tuned modifiers. hire, contact, inquiries and work with me match the phrasing people use when they are open to work.

Step 5 - add location for geography-bound hiring. Practitioners sometimes state a city or country, but many do not, so expect location filtering to reduce volume noticeably.

Step 6 - clean before you send. Filter out rows where possiblyTruncated is true, split rows with and without a username, then run the remaining Kaggle Email Scraper results through email verification.


Use cases for the Kaggle Email Scraper

Use caseWho runs itTypical keywords
Data science and ML recruitingTechnical recruiters and in-house sourcersdata scientist, machine learning engineer, kaggle master
Developer tools and MLOps prospectingFounders and sales teams at ML vendorsdeep learning, pytorch, feature engineering
Developer relations and community buildingDevRel and community managersnotebook, tutorial, open source
Research collaboration outreachLabs, universities and research groupsnlp researcher, computer vision, time series
Bootcamp and course enrolmentEducation providers and training platformsbeginner, learning data science, portfolio
Freelance and contract data workAgencies and consultanciesfreelance data scientist, analytics consultant
Conference and hackathon recruitingEvent organisers and sponsorscompetition, hackathon, benchmark
Talent-market researchAnalysts studying the ML labour marketany role keyword, varied by location

Kaggle is a demonstrated-skill community rather than a resume site, which is why sourcers value it, and why the Kaggle Email Scraper is best used as one input among several rather than a standalone list.

For academic contacts the ResearchGate Email Scraper and the Academia.edu Email Scraper are usually richer, and practitioners who publish courses often appear in the Coursera Email Scraper or the Pluralsight Email Scraper.


Why choose this Kaggle Email Scraper

Manual sourcing on Kaggle means paging through leaderboards, opening profiles, reading bios and copying addresses. The Kaggle Email Scraper automates that loop and returns a deduplicated dataset.

Every row is auditable. The query and keyword fields show exactly how a lead was found, which matters more here than elsewhere because it lets you separate profile-shaped queries from dataset-shaped ones.

The extraction layer is deliberately conservative. Junk filtering, boundary-correct domain matching and soft-wrap repair mean far fewer garbage rows to clean downstream.

Long Kaggle Email Scraper runs survive interruption. State persists in the key-value store and saves fire on Apify's PERSIST_STATE, MIGRATING and ABORTING events.


Worked example: a computer vision hiring pipeline

You are hiring three computer vision engineers and want a top-of-funnel list of 300 practitioners.

Set keywords to ["computer vision", "image segmentation", "kaggle master"], customDomains to ["@gmail.com"], countryCode to "US" and maxEmails to 300, leaving expandQueries on.

The Kaggle Email Scraper issues base, quoted and intitle: variants plus one query per modifier, paginating each independently and streaming leads as they are found.

Export the Kaggle Email Scraper dataset to CSV, drop possiblyTruncated rows, then sort by whether username is populated so your sourcers work the identifiable profiles first.


Limitations you should know before running

The Kaggle Email Scraper is deliberately transparent about its boundaries. Read these before planning a large data collection run.

Weaker identity resolution than creator platforms. Many indexed kaggle.com pages are datasets, competitions and notebooks rather than user profiles, so a fair share of Kaggle Email Scraper rows carry an email with no handle attached.

Only publicly indexed emails. The Kaggle Email Scraper returns addresses that already appear in Google's index. Practitioners who never published an address will not surface.

Google's per-query ceiling. A single query caps out at roughly 300 results. Query expansion exists to work around that ceiling, which is why it defaults to on.

possiblyTruncated. When Google's snippet ellipsis touches an address, this field is true and the email may be incomplete. Verify those rows before sending.

Apify GOOGLE_SERP proxy required. The Kaggle Email Scraper cannot run without Apify proxy credentials that include the GOOGLE_SERP group.

Free-plan cap. Running the Kaggle Email Scraper on a free Apify plan limits results to 100 emails per run. Paid plans are uncapped.

Handle availability. username and profileUrl are populated only when Google's result exposes a Kaggle handle. Rows without one still carry accountName and, sometimes, fullName. This is a Google limitation, not a bug.

No guaranteed volume. Kaggle Email Scraper yield varies with keywords, domains and location. No result count is promised.

The Kaggle Email Scraper is not affiliated with, endorsed by or officially supported by Kaggle.


Every Actor below runs the same engine as the Kaggle Email Scraper, with identical input fields and an identical output schema; only the target platform differs.

ActorWhat it collects
Kaggle Email and Phone Number ScraperEmails and phone numbers from Kaggle
Kaggle Phone Number ScraperPublic phone numbers from Kaggle
Academia.edu Email ScraperPublic contact emails from Academia.edu
Buy Me a Coffee Email ScraperPublic contact emails from Buy Me a Coffee
Carrd Email ScraperPublic contact emails from Carrd
Domestika Email ScraperPublic contact emails from Domestika
edX Email ScraperPublic contact emails from edX
Gumroad Email ScraperPublic contact emails from Gumroad
Kajabi Email ScraperPublic contact emails from Kajabi
Ko-fi Email ScraperPublic contact emails from Ko-fi
Linktree Email ScraperPublic contact emails from Linktree
MasterClass Email ScraperPublic contact emails from MasterClass
Meetup Email ScraperPublic contact emails from Meetup
Pluralsight Email ScraperPublic contact emails from Pluralsight
Preply Email ScraperPublic contact emails from Preply
ResearchGate Email ScraperPublic contact emails from ResearchGate
Skillshare Email ScraperPublic contact emails from Skillshare
Teachable Email ScraperPublic contact emails from Teachable
Udemy Email ScraperPublic contact emails from Udemy
Academia.edu Email and Phone Number ScraperEmails and phone numbers from Academia.edu
Buy Me a Coffee Email and Phone Number ScraperEmails and phone numbers from Buy Me a Coffee
Carrd Email and Phone Number ScraperEmails and phone numbers from Carrd
Domestika Email and Phone Number ScraperEmails and phone numbers from Domestika
edX Email and Phone Number ScraperEmails and phone numbers from edX
Gumroad Email and Phone Number ScraperEmails and phone numbers from Gumroad
Kajabi Email and Phone Number ScraperEmails and phone numbers from Kajabi
Ko-fi Email and Phone Number ScraperEmails and phone numbers from Ko-fi
Linktree Email and Phone Number ScraperEmails and phone numbers from Linktree
MasterClass Email and Phone Number ScraperEmails and phone numbers from MasterClass
Meetup Email and Phone Number ScraperEmails and phone numbers from Meetup
Pluralsight Email and Phone Number ScraperEmails and phone numbers from Pluralsight
Preply Email and Phone Number ScraperEmails and phone numbers from Preply
ResearchGate Email and Phone Number ScraperEmails and phone numbers from ResearchGate
Skillshare Email and Phone Number ScraperEmails and phone numbers from Skillshare
Teachable Email and Phone Number ScraperEmails and phone numbers from Teachable
Academia.edu Phone Number ScraperPublic phone numbers from Academia.edu
Buy Me a Coffee Phone Number ScraperPublic phone numbers from Buy Me a Coffee
Carrd Phone Number ScraperPublic phone numbers from Carrd
Domestika Phone Number ScraperPublic phone numbers from Domestika
edX Phone Number ScraperPublic phone numbers from edX
Gumroad Phone Number ScraperPublic phone numbers from Gumroad
Kajabi Phone Number ScraperPublic phone numbers from Kajabi

FAQ about the Kaggle Email Scraper

What is the Kaggle Email Scraper?

It is an Apify Actor that extracts publicly indexed contact emails from Kaggle by running domain-scoped Google searches and parsing the resulting search-result blocks.

Does the Kaggle Email Scraper log into Kaggle or use its API?

No. It never logs in, never calls a Kaggle API and never opens a browser, so there are no cookies and no authentication anywhere in the pipeline.

Why do so many rows have no username?

Because much of Kaggle's indexed surface is datasets, competitions and notebooks rather than profiles. Those pages have no handle for the Kaggle Email Scraper to resolve, so username is null and profileUrl is empty.

Can I improve the share of rows with a handle?

Often, yes. Role-shaped keywords such as data scientist, machine learning engineer or kaggle master return more profile pages than dataset-shaped keywords.

Where do the emails actually come from?

From publicly indexed titles, snippets and site labels, typically bios and descriptions where someone published an address themselves.

What does possiblyTruncated: true mean?

Google's snippet ellipsis touched the address, so it may be cut off. Verify those rows before contacting anyone.

How many emails will one run of the Kaggle Email Scraper return?

maxEmails accepts 1 to 10000, but free Apify plans are capped at 100 emails per run and paid plans are uncapped. Real volume depends on keywords, domains and location.

Is a proxy required?

Yes. The Kaggle Email Scraper requires the Apify GOOGLE_SERP proxy and cannot run without Apify proxy credentials.

Should I leave expandQueries on?

Almost always. Google returns roughly 300 results per query, and expansion multiplies your reachable surface through base, quoted, intitle: and modifier variants.

Can I target a specific country?

Yes. Set countryCode for the search proxy and optionally add a location phrase, keeping in mind that many practitioners never state a city.

What happens if my run is interrupted or migrated?

State persists in the key-value store keyed by a hash of your input, with throttled saves plus saves on PERSIST_STATE, MIGRATING and ABORTING, so an interrupted run resumes instead of restarting.

Can I email everyone in the dataset immediately?

Treat the Kaggle Email Scraper output as raw lead data. Deduplication is automatic, but verification, consent and compliance with GDPR, CAN-SPAM and Kaggle's terms remain your responsibility.


Leave a review

If the Kaggle Email Scraper saved you time, please leave a star rating and a short review on the Actor page.

Reviews are how other buyers judge whether a tool works, and they tell us which features to build next.

If something did not work, email neurodata.apify@gmail.com instead - bugs get fixed faster than they get complained about.

Support

Questions about the Kaggle Email Scraper, bug reports or a custom build? Email neurodata.apify@gmail.com.