edX Email Scraper
Under maintenancePricing
from $2.49 / 1,000 results
edX Email Scraper
Under maintenanceedX Email Scraper SD - edX Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from edX results by keyword, location and email domain - edX email extractor.
Pricing
from $2.49 / 1,000 results
Rating
0.0
(0)
Developer
Leads Scraper
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
12 days ago
Last modified
Categories
Share
edX Email Scraper — public emails from indexed course and institution pages
The edX Email Scraper searches publicly indexed edx.org pages with Google's site: operator and extracts contact emails from the titles and snippets it finds.
Everything runs through the Apify GOOGLE_SERP proxy. There is no login, no edX API, no browser and no cookies.
Understand the shape of the data first. edX publishes course and institution pages rather than personal instructor profiles. Most rows this Actor returns therefore carry a course or school name instead of a personal handle.
In practice that means accountName often holds an institution or programme label, fullName is empty, and both username and profileUrl are frequently null and "". The email and the snippet are still real.
That makes the edX Email Scraper a tool for institution-level and programme-level contact discovery — partner enquiry addresses, programme support lines, departmental contacts printed in a course description — rather than a directory of named instructors.
Quick start with the edX Email Scraper
- Open the edX Email Scraper on Apify and enter programme-style keywords such as
professional certificate data analysis. - Add every email domain you are willing to accept to
customDomains, including institutional ones. - Set
maxEmailsto a realistic ceiling and start the run. - Export the dataset as CSV, JSON or Excel once it finishes.
The edX Email Scraper streams rows into the dataset as it finds them, so you can tell within a minute whether a keyword is matching catalogue copy or real contact lines.
Expect institution-level rows: that is what the edX Email Scraper is good at, and it is what edX publishes.
What the edX Email Scraper actually does
The edX Email Scraper does not log into edX, does not use an edX API, and does not open the edX website. No browser, no JavaScript rendering, no authentication.
All data comes from publicly indexed Google search results: the result title, the site label and the snippet Google prints.
A generated query looks like site:edx.org professor department contact "@gmail.com". Result pages are fetched asynchronously with aiohttp through the search proxy.
Parsing is structural — the edX Email Scraper locates the <h3> title and then the smallest surrounding block, rather than depending on Google's CSS class names.
Key features of the edX Email Scraper
| Feature | What it means in practice |
|---|---|
Google site: targeting | Every query is scoped to edx.org |
| Query expansion | Base, quoted and intitle: phrasings plus one variant per modifier; base queries run first |
| Course-oriented modifiers | Defaults include enroll, coaching and workshop |
| Domain filtering | Only emails ending in one of your customDomains survive |
| Global deduplication | Each unique address is pushed to the dataset once |
| Email normalisation | Understands name [at] domain [dot] com, name (at) domain, name @ domain.com, domain .com, zero-width characters and the full-width @ |
| Junk filter | Rejects placeholders such as email@, yourname@, test@, xxx@ and single-character locals |
| Boundary-correct matching | @gmail.com will not match inside @gmail.company or @gmail.com.br |
| Soft-wrap repair | Drops a hit that is only the tail of another address in the same block |
| Async concurrency | asyncio worker pool with a shared stop signal on maxEmails |
| Retries | Up to 3 attempts per page, exponential backoff, fresh proxy session per request |
| Block detection | CAPTCHA, "unusual traffic" and consent pages are detected and retried, not treated as empty |
| Re-queue | Blocked or failed queries are retried once at the end of the run |
| Resumable state | Key-value store state keyed by a hash of the input; saves fire on PERSIST_STATE, MIGRATING and ABORTING |
| Whole-page fallback parser | A Google layout change degrades to "emails without account details", not "no emails" |
| Truncation flag | possiblyTruncated marks addresses touched by Google's snippet ellipsis |
| Run summary | Logs pages fetched, blocked pages, retries and emails per page |
How the edX Email Scraper works
- Reads the input: keywords, location, allowed email domains and run limits.
- Builds Google queries with
site:edx.org, one per keyword × domain × phrasing. - Fetches result pages through the Apify GOOGLE_SERP proxy.
- Locates each result block structurally and extracts the title, institution label and snippet.
- Applies the domain-filtered regex, normalisation and junk filtering to the block text.
- Deduplicates globally and pushes each lead straight to the dataset.
Because the edX Email Scraper writes rows as it finds them, you can inspect early results and stop the run if the keyword is producing course marketing copy rather than contact lines.
Input reference for the edX Email Scraper
| Field | Type | Default | Meaning |
|---|---|---|---|
keywords | array (required) | ["instructor", "professor"] | Search terms (subject, programme, job title) |
location | string | "" | Optional location phrase added to every query |
customDomains | array | ["@gmail.com","@yahoo.com"] | Only emails on these domains are kept; @ optional |
maxEmails | integer 1–10000 | 20 | Stop after this many unique emails |
countryCode | string | "" | Two-letter country for the search proxy (US, GB, DE…) |
expandQueries | boolean | true | Search each keyword × domain in several phrasings |
queryModifiers | array | ["email","contact","enroll","coaching","workshop"] | Extra words combined with each keyword when expansion is on |
maxPagesPerQuery | integer 1–50 | 30 | Page cap per query |
maxConcurrency | integer 1–20 | 5 | Parallel queries |
Example input
{"keywords": ["micromasters supply chain", "computer science professor", "professional certificate cybersecurity"],"location": "","customDomains": ["@gmail.com", "@yahoo.com", "@outlook.com"],"maxEmails": 60,"countryCode": "US","expandQueries": true,"queryModifiers": ["email", "contact", "enroll", "coaching", "workshop"],"maxPagesPerQuery": 30,"maxConcurrency": 5}
Programme names — MicroMasters, professional certificate, executive education — work better as edX Email Scraper keywords than bare job titles, because that is the language edX pages actually use.
Output fields from the edX Email Scraper
Every dataset item produced by the edX Email Scraper contains all fourteen fields below.
| Field | Meaning |
|---|---|
network | Platform name |
keyword | The keyword that produced the lead |
query | The exact Google query used |
title | Raw result title |
accountName | Account label Google prints — on edX usually a school, partner or programme name |
fullName | Display name parsed from a profile-style title; empty for course and institution titles |
username | URL-safe handle when edX exposes one; otherwise null |
profileUrl | Canonical account URL when a handle is known; otherwise empty |
url | Direct platform link when exposed, else the profile URL |
description | Snippet text, cleaned of labels and counters |
email | Lower-cased email address |
emailDomain | The matched domain, e.g. @gmail.com |
possiblyTruncated | true when Google's snippet ellipsis touched the email |
foundAt | ISO 8601 UTC timestamp |
Example output
The first row is the ordinary edX case — an institution-level address with no personal handle behind it. The second row shows the rarer case where a bio page did expose one.
[{"network": "edX","keyword": "micromasters supply chain","query": "site:edx.org micromasters supply chain contact \"@gmail.com\"","title": "Supply Chain Analytics MicroMasters Program | edX","accountName": "edX","fullName": "","username": null,"profileUrl": "","url": "","description": "... Programme admissions and cohort enquiries: scm.programme.office@gmail.com. Enrol now ...","email": "scm.programme.office@gmail.com","emailDomain": "@gmail.com","possiblyTruncated": false,"foundAt": "2026-08-31T13:15:52.310Z"},{"network": "edX","keyword": "computer science professor","query": "site:edx.org intitle:\"computer science professor\" email \"@yahoo.com\"","title": "Ravi Subramanian | edX","accountName": "ravi-subramanian","fullName": "Ravi Subramanian","username": "ravi-subramanian","profileUrl": "https://www.edx.org/bio/ravi-subramanian","url": "https://www.edx.org/bio/ravi-subramanian","description": "Lecturer in computer science. Course correspondence: r.subramanian.cs@yahoo.com","email": "r.subramanian.cs@yahoo.com","emailDomain": "@yahoo.com","possiblyTruncated": false,"foundAt": "2026-08-31T13:16:21.774Z"}]
Use cases for the edX Email Scraper
edX Email Scraper data is institution-shaped, so the use cases that work are the ones aimed at programmes and schools rather than named individuals.
| Use case | How the edX Email Scraper helps |
|---|---|
| University partner and programme outreach | Snippets often carry a programme office or admissions address |
| EdTech B2B prospecting | accountName and title identify the partner institution |
| Curriculum and competitor research | Course descriptions summarise positioning and level |
| Corporate learning and L&D sourcing | Find programmes matching a skills gap, then contact the listed address |
| Academic content partnerships | Reach the programme contact printed on the course page |
| Accreditation and credential research | Professional certificate pages name the awarding body |
| Regional programme mapping | Add a location phrase to bias toward institutions in a region |
| Multi-platform education list building | Merge edX Email Scraper rows with sibling Actors for coverage |
Where an address belongs to an academic and appears alongside their published course material, it was published so that learners and peers could write about that material. Keep your message relevant to it.
Every address here is personal or organisational data under GDPR and comparable laws. Have a lawful basis, identify yourself, respect institutional contact norms, and act on opt-outs immediately.
For creator-published contact details in higher volume, see the Udemy Email Scraper and the Teachable Email Scraper; for scholarly correspondence lines, the ResearchGate Email Scraper and the Academia.edu Email Scraper are the better sources.
Getting better runs from the edX Email Scraper
Use programme vocabulary in the edX Email Scraper. MicroMasters, professional certificate and executive education match how edX titles its pages; instructor does not.
Broaden customDomains early in the edX Email Scraper. Institutional pages often print a departmental address on a university domain rather than a consumer one, so add those domains explicitly if you want them.
Leave expandQueries on. Google caps one query at roughly 300 results, and the modifier variants are how the edX Email Scraper widens coverage.
Keep maxEmails realistic. The edX Email Scraper reads a course-catalogue index, not a contact directory, and a modest ceiling gives you a clean, quick run.
Limitations of the edX Email Scraper
- edX publishes course and institution pages rather than personal instructor profiles, so most rows carry a course or school name instead of a handle. Expect empty
fullName,username: nullandprofileUrl: ""on the majority of rows. - The edX Email Scraper only finds addresses that are publicly visible in Google's index.
- Google caps a single query at roughly 300 results, which is exactly why query expansion exists.
usernameandprofileUrlpopulate only when Google's result exposes a handle. This is a Google limitation, not a bug.possiblyTruncated: truemeans Google's snippet ellipsis touched the address — verify it before sending.- Requires the Apify GOOGLE_SERP proxy; the edX Email Scraper cannot run without Apify proxy credentials.
- Free Apify plans are capped at 100 emails per run. Paid plans are uncapped.
- Results vary with keywords, domains and location, and no volume is guaranteed.
edX Email Scraper FAQ
How many results should I expect from the edX Email Scraper?
No volume is guaranteed, and this is not a high-yield platform. edX pages are catalogue entries, so only a minority print a contact address at all. Run several programme-style keywords with a wide domain list and treat maxEmails as a ceiling rather than a forecast.
Why are username and profileUrl usually empty?
Because the indexed pages are courses and institution listings, not personal profiles. With no handle in Google's result there is nothing to build a profile URL from, so the edX Email Scraper writes null and an empty string instead of guessing.
Why does accountName show a school instead of a person?
That is the label Google prints for these pages, and the edX Email Scraper reports it verbatim. On edX the account-level identity is usually the partner institution or the programme, which is exactly what the field reports.
Does the edX Email Scraper log into edX?
No. It does not log in, use an edX API, or open the edX website. It reads Google search results only.
Can I collect university domain addresses?
Yes. Add the domains you want — for example a specific university domain — to the edX Email Scraper's customDomains list, and only matching addresses are kept.
What does possiblyTruncated mean?
The edX Email Scraper flags a row when Google's ellipsis cut the snippet where the address sits, so the captured value may be incomplete. Keep those rows aside and verify them.
Do I need Apify proxy credentials?
Yes. The GOOGLE_SERP proxy is required and the edX Email Scraper cannot run without it.
Why is expandQueries enabled by default?
Because one Google query returns roughly 300 results at most. Expansion adds quoted, intitle: and modifier phrasings so the edX Email Scraper reaches more of the index.
Will a run resume after a migration?
Yes. The edX Email Scraper keeps state in the key-value store keyed by a hash of your input, with saves on PERSIST_STATE, MIGRATING and ABORTING.
What if Google changes its markup?
The edX Email Scraper parses structurally with a whole-page fallback, so a layout change degrades the run to "emails without account details" rather than returning nothing.
Is using this compliant with privacy law?
The edX Email Scraper reads only publicly indexed pages, but the addresses are still personal data. Under GDPR and similar regimes you need a lawful basis, clear identification and a working opt-out.
Can I export the dataset?
Yes. The edX Email Scraper dataset exports to JSON, CSV, Excel and the other Apify formats, and is reachable through the Apify API.
Related Actors
| Actor | What it collects |
|---|---|
| edX Email and Phone Number Scraper | Emails and phone numbers from edX |
| edX Phone Number Scraper | Public phone numbers from edX |
| Academia.edu Email Scraper | Public contact emails from Academia.edu |
| Buy Me a Coffee Email Scraper | Public contact emails from Buy Me a Coffee |
| Carrd Email Scraper | Public contact emails from Carrd |
| Domestika Email Scraper | Public contact emails from Domestika |
| Gumroad Email Scraper | Public contact emails from Gumroad |
| Kaggle Email Scraper | Public contact emails from Kaggle |
| Kajabi Email Scraper | Public contact emails from Kajabi |
| Ko-fi Email Scraper | Public contact emails from Ko-fi |
| Linktree Email Scraper | Public contact emails from Linktree |
| MasterClass Email Scraper | Public contact emails from MasterClass |
| Meetup Email Scraper | Public contact emails from Meetup |
| Pluralsight Email Scraper | Public contact emails from Pluralsight |
| Preply Email Scraper | Public contact emails from Preply |
| ResearchGate Email Scraper | Public contact emails from ResearchGate |
| Skillshare Email Scraper | Public contact emails from Skillshare |
| Teachable Email Scraper | Public contact emails from Teachable |
| Udemy Email Scraper | Public contact emails from Udemy |
| Academia.edu Email and Phone Number Scraper | Emails and phone numbers from Academia.edu |
| Buy Me a Coffee Email and Phone Number Scraper | Emails and phone numbers from Buy Me a Coffee |
| Carrd Email and Phone Number Scraper | Emails and phone numbers from Carrd |
| Domestika Email and Phone Number Scraper | Emails and phone numbers from Domestika |
| Gumroad Email and Phone Number Scraper | Emails and phone numbers from Gumroad |
| Kaggle Email and Phone Number Scraper | Emails and phone numbers from Kaggle |
| Kajabi Email and Phone Number Scraper | Emails and phone numbers from Kajabi |
| Ko-fi Email and Phone Number Scraper | Emails and phone numbers from Ko-fi |
| Linktree Email and Phone Number Scraper | Emails and phone numbers from Linktree |
| MasterClass Email and Phone Number Scraper | Emails and phone numbers from MasterClass |
| Meetup Email and Phone Number Scraper | Emails and phone numbers from Meetup |
| Pluralsight Email and Phone Number Scraper | Emails and phone numbers from Pluralsight |
| Preply Email and Phone Number Scraper | Emails and phone numbers from Preply |
| ResearchGate Email and Phone Number Scraper | Emails and phone numbers from ResearchGate |
| Skillshare Email and Phone Number Scraper | Emails and phone numbers from Skillshare |
| Teachable Email and Phone Number Scraper | Emails and phone numbers from Teachable |
| Academia.edu Phone Number Scraper | Public phone numbers from Academia.edu |
| Buy Me a Coffee Phone Number Scraper | Public phone numbers from Buy Me a Coffee |
| Carrd Phone Number Scraper | Public phone numbers from Carrd |
| Domestika Phone Number Scraper | Public phone numbers from Domestika |
| Gumroad Phone Number Scraper | Public phone numbers from Gumroad |
| Kaggle Phone Number Scraper | Public phone numbers from Kaggle |
| Kajabi Phone Number Scraper | Public phone numbers from Kajabi |
Leave a review
If the edX Email Scraper saved you time, please leave a star rating and a short review on the Actor page.
Reviews are how other buyers judge whether a tool works, and they tell us which features to build next.
If something did not work, email neurodata.apify@gmail.com instead - bugs get fixed faster than they get complained about.
Support
Questions, bug reports or a custom build? Email neurodata.apify@gmail.com with your keywords, domain list and the result you were expecting.