Wikipedia Email Scraper - Domain Filter & Deduplication avatar

Wikipedia Email Scraper - Domain Filter & Deduplication

Pricing

from $0.99 / 1,000 results

Go to Apify Store
Wikipedia Email Scraper - Domain Filter & Deduplication

Wikipedia Email Scraper - Domain Filter & Deduplication

๐Ÿ“– Wikipedia Email Scraper โ€” turn Wikipedia searches into a clean contributor and organisation email list. Filter by keyword, location and domain, merge alias duplicates and decode hidden addresses. ๐Ÿ” For research & PR outreach.

Pricing

from $0.99 / 1,000 results

Rating

0.0

(0)

Developer

InsightFlow

InsightFlow

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

Wikipedia Email Scraper

Wikipedia Email Scraper helps you extract publicly available email addresses from Wikipedia pages using keywords and filters you choose. Itโ€™s built for Wikipedia scraping, Wikipedia data extraction, and contact discovery, making it a practical email scraper for lead generation, research, and structured data extraction at scale. ๐Ÿš€

What Does Wikipedia Email Scraper Do? ๐Ÿค–

Wikipedia Email Scraper searches public Wikipedia content based on your keywords, optional location, and selected email-domain filters. During a run, it scrapes relevant pages, extracts matching emails, filters duplicates, and exports the results to your Apify dataset in a clean, structured format. You can control the total number of emails with maxEmails, which helps with pacing and budget control. This is a faster, more scalable alternative to manual wiki crawler workflows and repetitive text mining.

What Can Wikipedia Email Scraper Extract? ๐Ÿ“Š

This actor captures a small but highly useful set of fields for contact scraping and lead generation. Each record includes the keyword that surfaced the result, the page title, a text description, the page URL, and the extracted email address. That makes it easy to review leads, analyze context, and export the results into your own workflow.

Data TypeField NameDescription
DiscoverykeywordThe keyword that led to this result
IdentitytitleThe page title shown in the dataset
ContextdescriptionText description associated with the page result
NavigationurlDirect link to the Wikipedia page
ContactemailExtracted email address from publicly available data

Key Features of Wikipedia Email Scraper โšก

  • โœ… Keyword-Driven Extraction: Uses your keywords to find relevant Wikipedia pages and collect matching contact data.
  • ๐ŸŒ Optional Location Filter: Narrow results by location when you want more focused Wikipedia scraping.
  • ๐Ÿ“ง Custom Domain Filtering: Limit email extraction to selected domains like @gmail.com or @yahoo.com.
  • ๐Ÿ”„ Reliable Run Behavior: Includes retries, fallbacks, and progress-saving logic for more resilient web scraping.
  • ๐Ÿ“Š Structured Dataset Output: Saves each result as a clean record for analysis, export, or CRM use.
  • ๐Ÿ’พ Live Saving: Results are pushed during the run, so you do not lose progress if the run stops early.
  • โš™๏ธ Configurable Volume: Set maxEmails to cap how many emails are collected per run.

How to Use Wikipedia Email Scraper ๐Ÿš€

  1. Open the Actor โ€” Find Wikipedia Email Scraper in the Apify Store.
  2. Add Keywords โ€” Enter one or more keywords such as job titles or roles.
  3. Set Optional Filters โ€” Use location and custom email domains to refine results.
  4. Choose a Limit โ€” Set maxEmails to control how many emails the actor collects.
  5. Run the Actor โ€” Start the run and monitor progress in the logs.
  6. Review Results โ€” Open the Dataset tab to inspect and export your scraped leads.

No coding required. โœจ

Wikipedia Email Scraper Output Format ๐Ÿ“ฆ

The actor saves results in the dataset using the exact fields shown below. This makes it easy to export the data for analysis, lead generation, or further structured data extraction.

Input Example

{
"keywords": [
"manager",
"founder"
],
"location": "",
"customDomains": [
"@gmail.com",
"@yahoo.com"
],
"maxEmails": 20
}
ParameterTypeRequiredDefaultDescription
keywordsArrayโœ… Yes["manager","founder"]One or more keywords or queries used to find relevant Wikipedia pages.
locationStringNo""Optional location filter to narrow down search results.
customDomainsArrayNo["@gmail.com","@yahoo.com"]Email domains to include when extracting contacts.
maxEmailsIntegerNo20Maximum number of emails to collect before the actor stops.

Output Example

[
{
"keyword": "manager",
"title": "Jane Smith",
"description": "Marketing manager and public speaker based in London.",
"url": "https://en.wikipedia.org/wiki/Jane_Smith",
"email": "jane.smith@gmail.com"
}
]
FieldLabelFormatDescription
keywordKeywordtextThe keyword used to find the result.
titleTitletextThe page title returned in the dataset.
descriptionDescriptiontextText context associated with the result.
urlUrllinkDirect link to the Wikipedia page.
emailEmailtextExtracted email address found in publicly available data.

๐ŸŽฏ Use Cases of Wikipedia Email Scraper

Lead Generation: Build targeted lists of publicly available contacts for outreach campaigns, prospecting, and contact discovery.

Email Marketing Campaigns: Collect email addresses from relevant Wikipedia pages to support newsletters, follow-ups, and promotional workflows.

Research and Entity Extraction: Use Wikipedia scraping for Wikipedia parser-style analysis, text mining, and structured data extraction across topics.

Market Research: Gather page titles, descriptions, and contact details to understand niches, industries, and public-facing entities.

CRM Enrichment: Add extracted email data to existing records to improve segmentation and improve contact completeness.

How Much Will Wikipedia Email Scraper Cost You? ๐Ÿ’ฐ

This actor is designed for controlled, practical runs with a configurable maxEmails limit, so you can manage output volume and keep costs predictable. Apify pricing depends on your platform usage, and setting a lower limit helps you keep each run focused. For larger campaigns, increase maxEmails carefully and review results in the Dataset tab as they arrive. ๐Ÿ“ˆ

Wikipedia Email Scraper only accesses publicly available data on Wikipedia. It does not require login access or private content, and it is intended for legitimate research, marketing, and contact discovery use cases. As with any web scraping or email harvesting workflow, you are responsible for complying with applicable laws, data-protection rules, and Wikipediaโ€™s terms. If you have questions, reach out at insightflowofficial@gmail.com.

Wikipedia Email Scraper Input Parameters ๐Ÿ“‹

{
"keywords": [
"manager",
"founder"
],
"location": "",
"customDomains": [
"@gmail.com",
"@yahoo.com"
],
"maxEmails": 20
}
ParameterTypeRequiredDefaultDescription
keywordsArrayโœ… Yes["manager","founder"]A list of keywords or queries to search for.
locationStringNo""Location to filter search results.
customDomainsArrayNo["@gmail.com","@yahoo.com"]List of custom email domains.
maxEmailsIntegerNo20Maximum number of emails to collect. The actor stops once this limit is reached.

During the Actor Run โฑ๏ธ

Youโ€™ll see live logs in the Apify Console as the actor progresses through your keywords and email-domain filters. Results are saved in real time to the dataset, so you can inspect leads before the run finishes. Runtime depends on your input size, especially maxEmails, keyword breadth, and how much publicly available data matches your filters. If results are limited, try broader keywords or additional domains.

Final Note โœ‰๏ธ

Start extracting Wikipedia emails in minutes with a simple, scalable Wikipedia Email Scraper. Itโ€™s an efficient way to automate Wikipedia data extraction and contact scraping with clean results you can use right away. Questions? Contact insightflowofficial@gmail.com.

FAQ โ€” Wikipedia Email Scraper โ“

How does the Wikipedia Email Scraper find emails?

It uses your keywords and email-domain filters to locate relevant Wikipedia pages, then extracts publicly available email addresses from the page content. The actor returns only matched contact data that is visible on the web.

What types of Wikipedia pages can I scrape?

You can scrape public Wikipedia pages that contain text and visible contact details matching your chosen keywords. If a page does not contain a relevant email address, it will not produce a contact result.

Why use Wikipedia scraping for lead generation?

Wikipedia scraping can support lead generation, research, and contact discovery when you want structured data from public web pages. It saves time by replacing manual review with automated extraction.

How much does the Wikipedia Email Scraper cost?

The actor supports controlled runs through the maxEmails parameter, so you can decide how much data to collect per run. That makes it easier to manage usage and keep scraping focused.

What are the main limitations of the Wikipedia Email Scraper?

Results depend on what publicly available data exists on the page. If a page has no matching email address, no result will be saved. Larger searches can also take longer, especially with broader keywords.

How do I choose the right keywords?

Use targeted keywords related to the type of contacts you want, then expand with similar terms if results are too limited. This usually improves email extraction, entity extraction, and overall lead quality.

How does the Wikipedia Email Scraper help my business?

It turns manual Wikipedia data extraction into an automated workflow, helping you collect leads faster, enrich records, and export structured data for campaigns or analysis.

Who do I contact for support or feedback?

For support, feedback, or custom feature requests, contact insightflowofficial@gmail.com.

๐Ÿ†˜ Support & Feedback

Found a bug or need help with Wikipedia Email Scraper?

  • ๐Ÿž Bug reports: Share what happened and which input you used
  • โœจ Feature requests: Send your ideas for new filters or output improvements
  • ๐Ÿ“ง Email: insightflowofficial@gmail.com

MX Lookup

Every address is checked at the DNS level: the actor resolves the mail domain's MX records and reports what it found.

FieldMeaning
mxFoundtrue when the domain publishes at least one mail server
mxHostThe lowest-preference (primary) mail server
mxRecordsEvery MX record found, in preference order
mxProviderWho runs the mail: Google Workspace, Microsoft 365, Zoho, Proton, ...
mxStatusfound, no_records, no_such_domain, timeout, error, or skipped

Inputs

  • MX Lookup - turn the check on or off (default: on).
  • Only keep contacts whose domain has an MX record - drop unreachable domains. Only a definite negative (no_records / no_such_domain) drops a contact; a timeout or resolver error is treated as unknown and the lead is kept.
  • MX lookup timeout (seconds) - per-domain DNS budget, 1-15s.

This is a domain-level check. It confirms the domain can receive mail; it does not verify that an individual mailbox exists.