Wikipedia Email Scraper - Domain Filter & Deduplication
Pricing
from $0.99 / 1,000 results
Wikipedia Email Scraper - Domain Filter & Deduplication
๐ Wikipedia Email Scraper โ turn Wikipedia searches into a clean contributor and organisation email list. Filter by keyword, location and domain, merge alias duplicates and decode hidden addresses. ๐ For research & PR outreach.
Pricing
from $0.99 / 1,000 results
Rating
0.0
(0)
Developer
InsightFlow
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Wikipedia Email Scraper
Wikipedia Email Scraper helps you extract publicly available email addresses from Wikipedia pages using keywords and filters you choose. Itโs built for Wikipedia scraping, Wikipedia data extraction, and contact discovery, making it a practical email scraper for lead generation, research, and structured data extraction at scale. ๐
What Does Wikipedia Email Scraper Do? ๐ค
Wikipedia Email Scraper searches public Wikipedia content based on your keywords, optional location, and selected email-domain filters. During a run, it scrapes relevant pages, extracts matching emails, filters duplicates, and exports the results to your Apify dataset in a clean, structured format. You can control the total number of emails with maxEmails, which helps with pacing and budget control. This is a faster, more scalable alternative to manual wiki crawler workflows and repetitive text mining.
What Can Wikipedia Email Scraper Extract? ๐
This actor captures a small but highly useful set of fields for contact scraping and lead generation. Each record includes the keyword that surfaced the result, the page title, a text description, the page URL, and the extracted email address. That makes it easy to review leads, analyze context, and export the results into your own workflow.
| Data Type | Field Name | Description |
|---|---|---|
| Discovery | keyword | The keyword that led to this result |
| Identity | title | The page title shown in the dataset |
| Context | description | Text description associated with the page result |
| Navigation | url | Direct link to the Wikipedia page |
| Contact | email | Extracted email address from publicly available data |
Key Features of Wikipedia Email Scraper โก
- โ Keyword-Driven Extraction: Uses your keywords to find relevant Wikipedia pages and collect matching contact data.
- ๐ Optional Location Filter: Narrow results by location when you want more focused Wikipedia scraping.
- ๐ง Custom Domain Filtering: Limit email extraction to selected domains like
@gmail.comor@yahoo.com. - ๐ Reliable Run Behavior: Includes retries, fallbacks, and progress-saving logic for more resilient web scraping.
- ๐ Structured Dataset Output: Saves each result as a clean record for analysis, export, or CRM use.
- ๐พ Live Saving: Results are pushed during the run, so you do not lose progress if the run stops early.
- โ๏ธ Configurable Volume: Set
maxEmailsto cap how many emails are collected per run.
How to Use Wikipedia Email Scraper ๐
- Open the Actor โ Find Wikipedia Email Scraper in the Apify Store.
- Add Keywords โ Enter one or more keywords such as job titles or roles.
- Set Optional Filters โ Use location and custom email domains to refine results.
- Choose a Limit โ Set
maxEmailsto control how many emails the actor collects. - Run the Actor โ Start the run and monitor progress in the logs.
- Review Results โ Open the Dataset tab to inspect and export your scraped leads.
No coding required. โจ
Wikipedia Email Scraper Output Format ๐ฆ
The actor saves results in the dataset using the exact fields shown below. This makes it easy to export the data for analysis, lead generation, or further structured data extraction.
Input Example
{"keywords": ["manager","founder"],"location": "","customDomains": ["@gmail.com","@yahoo.com"],"maxEmails": 20}
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
keywords | Array | โ Yes | ["manager","founder"] | One or more keywords or queries used to find relevant Wikipedia pages. |
location | String | No | "" | Optional location filter to narrow down search results. |
customDomains | Array | No | ["@gmail.com","@yahoo.com"] | Email domains to include when extracting contacts. |
maxEmails | Integer | No | 20 | Maximum number of emails to collect before the actor stops. |
Output Example
[{"keyword": "manager","title": "Jane Smith","description": "Marketing manager and public speaker based in London.","url": "https://en.wikipedia.org/wiki/Jane_Smith","email": "jane.smith@gmail.com"}]
| Field | Label | Format | Description |
|---|---|---|---|
keyword | Keyword | text | The keyword used to find the result. |
title | Title | text | The page title returned in the dataset. |
description | Description | text | Text context associated with the result. |
url | Url | link | Direct link to the Wikipedia page. |
email | text | Extracted email address found in publicly available data. |
๐ฏ Use Cases of Wikipedia Email Scraper
Lead Generation: Build targeted lists of publicly available contacts for outreach campaigns, prospecting, and contact discovery.
Email Marketing Campaigns: Collect email addresses from relevant Wikipedia pages to support newsletters, follow-ups, and promotional workflows.
Research and Entity Extraction: Use Wikipedia scraping for Wikipedia parser-style analysis, text mining, and structured data extraction across topics.
Market Research: Gather page titles, descriptions, and contact details to understand niches, industries, and public-facing entities.
CRM Enrichment: Add extracted email data to existing records to improve segmentation and improve contact completeness.
How Much Will Wikipedia Email Scraper Cost You? ๐ฐ
This actor is designed for controlled, practical runs with a configurable maxEmails limit, so you can manage output volume and keep costs predictable. Apify pricing depends on your platform usage, and setting a lower limit helps you keep each run focused. For larger campaigns, increase maxEmails carefully and review results in the Dataset tab as they arrive. ๐
Is It Legal to Scrape Wikipedia? โ๏ธ
Wikipedia Email Scraper only accesses publicly available data on Wikipedia. It does not require login access or private content, and it is intended for legitimate research, marketing, and contact discovery use cases. As with any web scraping or email harvesting workflow, you are responsible for complying with applicable laws, data-protection rules, and Wikipediaโs terms. If you have questions, reach out at insightflowofficial@gmail.com.
Wikipedia Email Scraper Input Parameters ๐
{"keywords": ["manager","founder"],"location": "","customDomains": ["@gmail.com","@yahoo.com"],"maxEmails": 20}
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
keywords | Array | โ Yes | ["manager","founder"] | A list of keywords or queries to search for. |
location | String | No | "" | Location to filter search results. |
customDomains | Array | No | ["@gmail.com","@yahoo.com"] | List of custom email domains. |
maxEmails | Integer | No | 20 | Maximum number of emails to collect. The actor stops once this limit is reached. |
During the Actor Run โฑ๏ธ
Youโll see live logs in the Apify Console as the actor progresses through your keywords and email-domain filters. Results are saved in real time to the dataset, so you can inspect leads before the run finishes. Runtime depends on your input size, especially maxEmails, keyword breadth, and how much publicly available data matches your filters. If results are limited, try broader keywords or additional domains.
Final Note โ๏ธ
Start extracting Wikipedia emails in minutes with a simple, scalable Wikipedia Email Scraper. Itโs an efficient way to automate Wikipedia data extraction and contact scraping with clean results you can use right away. Questions? Contact insightflowofficial@gmail.com.
FAQ โ Wikipedia Email Scraper โ
How does the Wikipedia Email Scraper find emails?
It uses your keywords and email-domain filters to locate relevant Wikipedia pages, then extracts publicly available email addresses from the page content. The actor returns only matched contact data that is visible on the web.
What types of Wikipedia pages can I scrape?
You can scrape public Wikipedia pages that contain text and visible contact details matching your chosen keywords. If a page does not contain a relevant email address, it will not produce a contact result.
Why use Wikipedia scraping for lead generation?
Wikipedia scraping can support lead generation, research, and contact discovery when you want structured data from public web pages. It saves time by replacing manual review with automated extraction.
How much does the Wikipedia Email Scraper cost?
The actor supports controlled runs through the maxEmails parameter, so you can decide how much data to collect per run. That makes it easier to manage usage and keep scraping focused.
What are the main limitations of the Wikipedia Email Scraper?
Results depend on what publicly available data exists on the page. If a page has no matching email address, no result will be saved. Larger searches can also take longer, especially with broader keywords.
How do I choose the right keywords?
Use targeted keywords related to the type of contacts you want, then expand with similar terms if results are too limited. This usually improves email extraction, entity extraction, and overall lead quality.
How does the Wikipedia Email Scraper help my business?
It turns manual Wikipedia data extraction into an automated workflow, helping you collect leads faster, enrich records, and export structured data for campaigns or analysis.
Who do I contact for support or feedback?
For support, feedback, or custom feature requests, contact insightflowofficial@gmail.com.
๐ Support & Feedback
Found a bug or need help with Wikipedia Email Scraper?
- ๐ Bug reports: Share what happened and which input you used
- โจ Feature requests: Send your ideas for new filters or output improvements
- ๐ง Email: insightflowofficial@gmail.com
MX Lookup
Every address is checked at the DNS level: the actor resolves the mail domain's MX records and reports what it found.
| Field | Meaning |
|---|---|
mxFound | true when the domain publishes at least one mail server |
mxHost | The lowest-preference (primary) mail server |
mxRecords | Every MX record found, in preference order |
mxProvider | Who runs the mail: Google Workspace, Microsoft 365, Zoho, Proton, ... |
mxStatus | found, no_records, no_such_domain, timeout, error, or skipped |
Inputs
- MX Lookup - turn the check on or off (default: on).
- Only keep contacts whose domain has an MX record - drop unreachable domains.
Only a definite negative (
no_records/no_such_domain) drops a contact; a timeout or resolver error is treated as unknown and the lead is kept. - MX lookup timeout (seconds) - per-domain DNS budget, 1-15s.
This is a domain-level check. It confirms the domain can receive mail; it does not verify that an individual mailbox exists.