# Researcher Email Finder: PubMed Authors by Topic (`datagleaner/researcher-email-finder`) Actor

Find researchers by topic with the email they published in PubMed papers, plus affiliation, ORCID and recent papers. $0.005 per researcher with email. Corresponding-author emails for life-science sales, CROs, recruiters and AI agents.

- **URL**: https://apify.com/datagleaner/researcher-email-finder.md
- **Developed by:** [Data Gleaner](https://apify.com/datagleaner) (community)
- **Categories:** Lead generation, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$5.00 / 1,000 researcher with emails

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Researcher Email Finder: PubMed Authors by Topic

**Find researchers by topic with the email they published in PubMed papers, plus affiliation, ORCID and recent papers. $0.005 per researcher with email.** Answers "corresponding author email", "pubmed author email", "pubmed emails", "researcher contact email" and "academic email finder" for any research field: give it a topic such as `crispr base editing`, an author name, or a list of PMIDs or DOIs, and get **one row per researcher** with the address they printed in a paper for correspondence, their institution, country, ORCID iD and their latest papers on the topic. No login, no API key, no CAPTCHA.

Built to be called by AI agents ("find researchers working on CRISPR base editing with their emails") and by sales and recruiting pipelines as well as people: the same fields on every row, deduplicated by email, and the price tied to the answer you want.

### Use cases

- **Life-science and lab-supply sales:** reach the labs publishing on the technique your reagent, instrument or software serves, with the address each corresponding author published.
- **CROs and biotech business development:** build a list of active investigators in an indication or modality ("car-t solid tumors", "spatial transcriptomics"), filtered by country or institution.
- **Conference, journal and webinar marketing:** invite the authors who published on your session topics in the last two years.
- **Academic and industry recruiting:** find postdocs and PIs publishing in a field, with their affiliation and ORCID.
- **AI agents and research assistants:** "who works on X, and how do I contact them?" as one tool call that charges only for researchers with an email.
- **Enrichment:** turn a list of PMIDs or DOIs into the contact emails of their corresponding authors.

### What it finds

PubMed records the affiliation of every author, and when a paper names a corresponding author, PubMed usually keeps the email printed with it (often as "Electronic address: ..."). This Actor searches PubMed for your topic, reads the full records of the most relevant papers through NCBI's official E-utilities API, and pulls out every author who published an email. Measured on test runs, **recent papers yield about 3 published emails for every 4 papers**, so 200 papers give roughly 150 researchers.

Each email is cleaned (trailing periods and "Electronic address:" labels removed, lower-cased) and given to the right author. When an affiliation prints several addresses, or several authors share one affiliation string, the address goes to the author whose name matches its local part (`wangfei@` goes to Fei Wang, not to his co-author Yang Wang); an address that cannot be tied to one person is dropped rather than guessed.

Researchers are deduplicated by email across all topics and papers in the run, and their other papers in the result (including ones where they did not print the email) are attached to the same row.

### How it works

1. Each topic or author name is searched on PubMed (relevance order), with `yearFrom` limiting it to recent papers.
2. The matching records are fetched in batches of up to 200 and parsed from PubMed's XML.
3. Emails are extracted from author affiliations and assigned to authors; institution and country are read from the affiliation, ORCID from the author record.
4. A search stops early once it alone has enough researchers for `maxResearchers`, or at `maxPapersScanned`.
5. Rows are spread evenly across your topics, newest-publishing researchers first, and capped at `maxResearchers`.

Requests are paced to NCBI's public limit (about 3 a second) and identify the tool to NCBI, as its usage policy asks.

### Input

| Field | Meaning |
|---|---|
| `queries` | Research topics, one per line. PubMed syntax works: `"base editing"[tiab] AND mice`. |
| `authorNames` | A named researcher: `Jennifer Doudna`, `Doudna JA` or `Doudna, Jennifer`. Only that author is returned. |
| `pmids`, `dois` | Specific papers. Every author on them who published an email is returned. |
| `yearFrom` | Only papers from this year on, e.g. `2024`. Applies to topics and author names. |
| `affiliationContains` | Keep researchers whose affiliations include this text (`Harvard`, `Hospital`). |
| `countries` | Keep researchers whose affiliation country is one of these (`United States`, `DE`, `UK`). |
| `onlyWithEmail` | On by default. Off also returns co-authors without an email, free. |
| `maxResearchers` | Rows to return across all topics (default 50). |
| `maxPapersScanned` | Papers read per topic or author (default 500). |
| `source` | `pubmed` (default, most emails) or `europepmc` (Europe PMC: ORCID for more authors, preprints, fewer emails). |
| `delayMs`, `proxyConfiguration` | Pacing (default 400 ms) and an optional proxy. No proxy by default. |

Leave every source field empty to run a built-in example: 5 researchers on `crispr base editing` (about 5 seconds, 5 x $0.005).

```json
{
  "queries": ["crispr base editing", "car-t cell therapy solid tumors", "spatial transcriptomics"],
  "yearFrom": 2024,
  "countries": ["United States", "United Kingdom", "Germany"],
  "maxResearchers": 100
}
```

### Output

One row per researcher, deduplicated by email. Fields that were not found are null, never invented.

```json
{
  "name": "Thomas Gaj",
  "firstName": "Thomas",
  "lastName": "Gaj",
  "email": "gaj@illinois.edu",
  "emailDomain": "illinois.edu",
  "isAcademicEmail": true,
  "hasEmail": true,
  "affiliation": "Department of Bioengineering, The Grainger College of Engineering, University of Illinois Urbana-Champaign, Urbana, IL, USA",
  "institution": "University of Illinois Urbana-Champaign",
  "country": "United States",
  "countryCode": "US",
  "orcid": "0000-0001-6004-9664",
  "orcidUrl": "https://orcid.org/0000-0001-6004-9664",
  "papers": [
    {
      "title": "In vivo CRISPR base editing for treatment of Huntington's disease",
      "pmid": "42527584",
      "doi": "10.1038/s41551-026-01747-y",
      "journal": "Nature biomedical engineering",
      "year": 2026,
      "url": "https://pubmed.ncbi.nlm.nih.gov/42527584/",
      "isCorresponding": true
    }
  ],
  "papersInResult": 11,
  "lastPublishedYear": 2026,
  "matchedQuery": "crispr base editing",
  "input": "query:crispr base editing",
  "scrapedAt": "2026-10-09T15:31:41+00:00"
}
```

- `papers` holds the 5 most recent papers in the result; `papersInResult` counts all of them. `isCorresponding` is true on papers where the researcher printed their email.
- `isAcademicEmail` is true for university, research-institute and public-research mailboxes, and false for free mail (Gmail, 163.com) and company domains.
- `input` says what found the row: `query:<topic>`, `author:<name>`, `pmid:<id>` or `doi:<doi>`.

The dataset has two views: **Researchers** (one row each) and **Papers** (one row per researcher and paper).

### Pricing

One event, pay per result:

- **`researcher-with-email`: $0.005** for each researcher returned with a published email.
- With `onlyWithEmail` on (default), researchers without an email are never returned. Turn it off and they are returned **free**.
- PubMed papers scanned are never charged; you pay for researchers, not for searching.

**Worked example.** Three topics with `maxResearchers: 150` return 150 researchers with emails: 150 x $0.005 = **$0.75**. In our test that run read 793 papers in 16 seconds.

Set a maximum total charge on the run and the Actor stops cleanly when it is reached.

### What to expect

- **Yield:** on recent papers (2024 on) in life-science topics, about 1.4 papers per email found with PubMed. Europe PMC prints fewer emails (about 2 papers per email in tests) but carries ORCID for more authors.
- **Who has an email:** usually the corresponding author, sometimes two or three per paper. First and middle authors rarely publish one.
- **Free-mail addresses:** some researchers publish a Gmail, 163.com or QQ address; `isAcademicEmail` lets you filter them.
- **Speed:** a few seconds per 200 papers.

### Limits

- Only emails the authors printed in the paper are returned. The Actor does not guess addresses from name patterns or look them up elsewhere.
- Emails can go stale when a researcher moves; `lastPublishedYear` and `yearFrom` help you keep to current addresses.
- Institution and country are read from the affiliation text, which journals format in many ways; a few rows have no country.
- PubMed covers biomedicine and life sciences. For other fields try `source: europepmc`, which also indexes preprints.
- Very common author names (`Wang Y`) match many people; give a full first name.

### Use with Python

```python
## pip install apify-client
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("datagleaner/researcher-email-finder").call(run_input={
    "queries": ["crispr base editing"],
    "yearFrom": 2024,
    "maxResearchers": 20,  # at most 20 x $0.005 = $0.10
})
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item["name"], item["email"], item["institution"], item["country"])
```

### Use with JavaScript / Node.js

```js
// npm install apify-client
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('datagleaner/researcher-email-finder').call({
    queries: ['spatial transcriptomics'],
    countries: ['Germany'],
    yearFrom: 2024,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const item of items) console.log(item.name, item.email, item.orcidUrl);
```

### Use it from n8n, Make, Zapier or an AI agent

Actor ID: `datagleaner/researcher-email-finder`

Minimal input:

```
{"queries": ["crispr base editing"], "yearFrom": 2024}
```

Each tool below runs this Actor with your own Apify API token.

- **n8n:** add the **Apify** node (`@apify/n8n-nodes-apify`). On n8n Cloud you install it from the community node registry. Choose **Run an Actor and get dataset**, set Actor to `datagleaner/researcher-email-finder` and paste the input above.

- **Make:** use the Apify app's **Run an Actor** module, then **Get Dataset Items** to read the results. **Watch Actor Runs** can trigger a scenario when a run finishes.

- **Zapier:** use the Apify action **Run Actor**, then the search **Fetch dataset items**. The trigger **Finished Actor run** starts a Zap when a run ends.

- **AI agents (MCP):** connect to `https://mcp.apify.com/?tools=datagleaner/researcher-email-finder`. In Claude Code:

  ```
  claude mcp add --transport http apify "https://mcp.apify.com/?tools=datagleaner/researcher-email-finder"
  ```

  Then run `/mcp` to sign in to Apify in your browser. Other clients can sign in with OAuth or send the header `Authorization: Bearer YOUR_APIFY_TOKEN`. Clients that run local MCP servers can use Apify's package (`@apify/actors-mcp-server`, run with `npx -y` and `APIFY_TOKEN` set) instead. Then ask the agent in plain words, for example:

  > Find 30 researchers in the US and UK who published on CRISPR base editing since 2024, with their emails, institutions and latest paper, as a table.

- **LangChain (Python):**

```python
## pip install langchain-apify, then set APIFY_TOKEN in your environment
import json
from langchain_apify import ApifyActorsTool
tool = ApifyActorsTool("datagleaner/researcher-email-finder")
result = tool.invoke({"run_input": json.loads('{"queries": ["crispr base editing"], "yearFrom": 2024}')})
```

### FAQ

**How do I find a corresponding author's email?** Give the paper's PMID or DOI, or a topic, to this Actor. It reads the paper's PubMed record and returns the email the corresponding author printed, with the author's affiliation.

**Where do the emails come from?** From the author affiliations in PubMed (or Europe PMC) records, where journals publish the corresponding author's address. Nothing is guessed or bought.

**Can I find a specific researcher's email?** Yes: put their name in `authorNames`, e.g. `Doudna JA`. If they published an email on any of their papers in range, it is returned with their papers.

**Does it get ORCID iDs?** Yes, when the record carries one. Europe PMC (`source: europepmc`) carries ORCID for more authors.

**Do I need an NCBI API key?** No. The Actor uses NCBI's public E-utilities within the published rate limit.

**Is it only for biomedicine?** PubMed is biomedical and life-science literature. Europe PMC adds preprints and some other fields.

### Responsible use

The emails this Actor returns are addresses the authors themselves published in their papers for scholarly correspondence. They are still personal data. You are responsible for lawful use, including GDPR, CAN-SPAM and local law (a lawful basis, an opt-out in every message, honouring removals) and for the terms of PubMed and Europe PMC. Do not use it for bulk unsolicited email: write to researchers individually, about their work, and only when your message is relevant to it.

# Actor input Schema

## `queries` (type: `array`):

Research topics to find researchers for, one per line, e.g. `crispr base editing` or `car-t solid tumors`. PubMed query syntax works: `"base editing"[tiab] AND mice`. Each topic is searched separately. Leave this and the other sources empty to run a 5-researcher example.

## `authorNames` (type: `array`):

Find a named researcher's published email: `Jennifer Doudna`, `Doudna JA` or `Doudna, Jennifer`. Only that author is returned from their papers.

## `pmids` (type: `array`):

Specific papers by PMID, e.g. `42850393`. Every author of these papers who printed an email is returned.

## `dois` (type: `array`):

Specific papers by DOI, e.g. `10.1038/s41551-026-01747-y` or a doi.org link.

## `yearFrom` (type: `integer`):

Only scan papers published in this year or later, so you reach researchers who are active now. Applies to topics and author names.

## `affiliationContains` (type: `string`):

Keep only researchers whose affiliation includes this text, e.g. `Harvard`, `Karolinska` or `Hospital`. Case-insensitive.

## `countries` (type: `array`):

Keep only researchers whose affiliation is in one of these countries, by name or ISO code: `United States`, `DE`, `UK`, `China`.

## `onlyWithEmail` (type: `boolean`):

On (default): return only researchers who published an email, the billed result. Off: also return co-authors without an email, free of charge, which can be many rows.

## `maxResearchers` (type: `integer`):

Stop after this many researchers across all topics. Results are spread evenly over the topics.

## `maxPapersScanned` (type: `integer`):

How many of the most relevant papers to read per topic or author. Roughly one paper in two to three yields a published email.

## `source` (type: `string`):

PubMed returns the most published emails. Europe PMC carries ORCID iDs for more authors and covers preprints, but prints fewer emails.

## `delayMs` (type: `integer`):

Minimum pause between requests. NCBI allows about 3 requests a second without an API key, so it stays at 340 ms or more.

## `proxyConfiguration` (type: `object`):

Optional. Defaults to no proxy: PubMed and Europe PMC are public APIs.

## Actor input object example

```json
{
  "queries": [
    "crispr base editing"
  ],
  "yearFrom": 2024,
  "onlyWithEmail": true,
  "maxResearchers": 50,
  "maxPapersScanned": 500,
  "source": "pubmed",
  "delayMs": 400
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "crispr base editing"
    ],
    "yearFrom": 2024
};

// Run the Actor and wait for it to finish
const run = await client.actor("datagleaner/researcher-email-finder").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["crispr base editing"],
    "yearFrom": 2024,
}

# Run the Actor and wait for it to finish
run = client.actor("datagleaner/researcher-email-finder").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "crispr base editing"
  ],
  "yearFrom": 2024
}' |
apify call datagleaner/researcher-email-finder --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datagleaner/researcher-email-finder"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/zeAaI807UDyfwCkSi/builds/n50d0mLVe6JaCOORk/openapi.json
