# PubMed Scraper (`aurenic/pubmed-scraper`) Actor

Extract biomedical articles from PubMed via the NCBI E-utilities API. 37M+ records: PMID, DOI, title, abstract, authors, journal, MeSH terms, keywords, references. No API key, no browser, no proxy.

- **URL**: https://apify.com/aurenic/pubmed-scraper.md
- **Developed by:** [Aurenic](https://apify.com/aurenic) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.20 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PubMed Scraper

Extract biomedical articles from PubMed via the NCBI E-utilities API. 37M+ records: PMID, DOI, title, abstract, authors, journal, MeSH terms, keywords, references. No API key, no browser, no proxy.

### What does PubMed Scraper do?

Scrape PubMed, the U.S. National Library of Medicine's biomedical literature database, in two modes:

- **Search by query** — full PubMed syntax: field tags (`cancer[Title]`), Boolean operators (AND/OR/NOT), quoted phrases, date filters (`2024[PDAT]`), and publication-type filters. Auto-paginates through results.
- **Fetch by PMID** — pass exact PubMed IDs, get full records in batch.

Every result includes the complete article metadata: PMID, DOI, PMC ID, title, journal, volume/issue/pages, publication date, author list with affiliations, full abstract (with labeled sections like "BACKGROUND:", "METHODS:"), MeSH descriptors with qualifiers, keywords, publication types, and reference count.

PubMed is fully public. **No API key required** — the actor sends the NCBI-required `tool` and `email` parameters with every request to stay in good standing.

### Output fields

| Field | Description |
|---|---|
| pmid | PubMed ID |
| doi | DOI identifier |
| pmcId | PubMed Central ID (open-access articles) |
| title | Article title |
| journal | Full journal name |
| journalAbbreviation | ISO journal abbreviation |
| volume / issue / pages / issn | Bibliographic details |
| publicationDate | Publication date |
| authors | Array of `{ lastName, firstName, initials, fullName, affiliation }` |
| authorCount / firstAuthor / lastAuthor | Author summary |
| abstract | Full abstract with section labels |
| abstractLength | Character count |
| meshTerms | Array of `{ descriptor, qualifiers[] }` |
| meshCount | Number of MeSH descriptors |
| keywords | Author-supplied keywords |
| publicationTypes | Review, Clinical Trial, Meta-Analysis, etc. |
| referenceCount | Number of cited references |
| language | Article language |
| pubmedUrl | Direct PubMed URL |

### Who is it for?

- **Systematic review teams** building PRISMA-compliant article datasets
- **Medical and life-science researchers** sourcing literature for meta-analyses
- **AI/ML researchers** building biomedical NLP corpora and training sets
- **Competitive intelligence in pharma and biotech** tracking publication activity
- **Medical writers and editors** finding evidence for regulatory submissions
- **Academic libraries** enriching catalogs with abstracts and MeSH data

### Pricing

**$1.20 per 1,000 results.** No subscription.

| Results | Cost |
|---|---|
| 100 | $0.12 |
| 1,000 | $1.20 |
| 10,000 | $12.00 |

### How to use it

1. Pick a **Mode**.
2. For search: enter a **Search Query** — PubMed syntax supported.
3. For fetch: enter **PMIDs**.
4. Optionally filter by **Date Range** and **Publication Types**.
5. Optionally set your real **Contact Email** (required by NCBI).
6. Click **Start**.

### Output example

```json
{
  "recordType": "article",
  "pmid": "38000000",
  "doi": "10.1038/s41586-023-06775-1",
  "pmcId": "PMC10705885",
  "title": "CRISPR-Cas9 genome editing in human primary T cells",
  "journal": "Nature",
  "journalAbbreviation": "Nature",
  "volume": "624",
  "issue": "7992",
  "pages": "415-423",
  "issn": "0028-0836",
  "publicationDate": "2023-12-14",
  "authors": [
    { "lastName": "Smith", "firstName": "Jane", "initials": "J", "fullName": "Jane Smith", "affiliation": "Department of Immunology, MIT" },
    { "lastName": "Chen", "firstName": "Wei", "initials": "W", "fullName": "Wei Chen", "affiliation": "Broad Institute" }
  ],
  "authorCount": 12,
  "firstAuthor": "Jane Smith",
  "lastAuthor": "David Liu",
  "abstract": "BACKGROUND: CRISPR-Cas9 has revolutionized genome editing...\n\nMETHODS: We engineered primary human T cells...\n\nRESULTS: Editing efficiency exceeded 95%...",
  "abstractLength": 1420,
  "meshTerms": [
    { "descriptor": "Gene Editing", "qualifiers": ["methods"] },
    { "descriptor": "CRISPR-Cas Systems", "qualifiers": [] },
    { "descriptor": "T-Lymphocytes", "qualifiers": ["metabolism"] }
  ],
  "meshCount": 14,
  "keywords": ["genome editing", "T cells", "CRISPR"],
  "publicationTypes": ["Journal Article", "Research Support, N.I.H., Extramural"],
  "referenceCount": 48,
  "language": "eng",
  "pubmedUrl": "https://pubmed.ncbi.nlm.nih.gov/38000000/",
  "scrapedAt": "2026-09-25T12:00:00.000Z"
}
```

### Technical details

- **Source: NCBI E-utilities API** at `https://eutils.ncbi.nlm.nih.gov/entrez/eutils/`. Free, no key, no login .
- **Rate limits: 3 req/s anonymous, 10 req/s with a free NCBI API key.** The actor defaults to 400ms between requests and backs off on 429 .
- **NCBI policy compliance** — every request sends `tool=` and `email=` parameters. This is NCBI's documented requirement to avoid IP blocks and to receive early warnings about policy changes.
- **esearch → efetch pipeline** — esearch returns PMIDs, efetch returns full XML records in batches of 200.
- **XML parsing** via `cheerio` in xmlMode — extracts structured records from the PubMed XML schema.
- **No browser, no proxy** — pure HTTP.

### Known limits

- **Rate limit is 3 req/s anonymous.** For high-volume runs, register a free NCBI API key to raise it to 10 req/s .
- **NCBI may block datacenter IPs** that overload the servers. The actor's `tool` and `email` params satisfy NCBI's registration requirement and prevent this — but running at the rate limit is not advised.
- **`retmax` caps at 10,000 PMIDs per search.** For deeper results, paginate with `retstart`.
- **efetch batches cap at ~200 PMIDs per request.** The actor handles this automatically.
- **Full text is not included.** PubMed provides abstracts and metadata; full text lives in PubMed Central (PMC) or publisher sites.
- **Retractions and updates** are noted in the PubMed record but not surfaced as separate fields.

### FAQ

**Do I need an API key?** No. Anonymous access works at 3 req/s. A free NCBI API key raises the limit to 10 req/s .

**Do I need a proxy?** No. Datacenter IPs work when sending the required `tool` and `email` params.

**What query syntax is supported?** Full PubMed syntax: field tags (`[Title]`, `[Author]`, `[PDAT]`), Boolean operators, quoted phrases, and filters. See the [PubMed search guide](https://pubmed.ncbi.nlm.nih.gov/help/).

**How do I search by MeSH term?** Use the `[MeSH Terms]` tag: `diabetes[MeSH Terms] AND 2024[PDAT]`.

**How do I search a specific author?** Use the `[Author]` tag: `Smith J[Author] AND cancer[Title]`.

**How do I export data?** After a run, go to Storage → Export as JSON, CSV, Excel.

### Support

Open an issue on the Actor's page for bugs or feature requests.

# Actor input Schema

## `mode` (type: `string`):

What to fetch.

## `query` (type: `string`):

PubMed query. Supports full PubMed syntax: field tags (e.g. 'cancer\[Title]'), Boolean operators (AND/OR/NOT), and quoted phrases.

## `pmids` (type: `array`):

PubMed IDs to fetch directly (e.g. 38000000). Used in fetch mode.

## `retmax` (type: `integer`):

PMIDs to retrieve per esearch call. Max 10,000.

## `retstart` (type: `integer`):

First result index for pagination.

## `sort` (type: `string`):

Result order.

## `dateFrom` (type: `string`):

Earliest publication date (YYYY/MM/DD). Applied as a PDAT filter.

## `dateTo` (type: `string`):

Latest publication date (YYYY/MM/DD).

## `publicationTypes` (type: `array`):

Filter by publication type: 'Review', 'Clinical Trial', 'Meta-Analysis', 'Randomized Controlled Trial', 'Systematic Review', 'Case Reports', 'Editorial', 'Letter'.

## `includeAbstract` (type: `boolean`):

Attach full abstract text to each record.

## `includeMeSH` (type: `boolean`):

Attach MeSH (Medical Subject Headings) descriptor and qualifier terms.

## `email` (type: `string`):

Required by NCBI policy for all E-utilities requests. Used for IP-block notices — should be a real address.

## `tool` (type: `string`):

Tool identifier sent with every request. Required by NCBI policy.

## `apiKey` (type: `string`):

Optional. Raises rate limit from 3 req/s to 10 req/s. Get a free key from your NCBI account at https://account.ncbi.nlm.nih.gov/.

## `maxItems` (type: `integer`):

Hard cap on articles per run.

## `requestDelayMs` (type: `integer`):

Delay between requests. NCBI allows 3 req/s without a key. Default 400ms stays safely under.

## Actor input object example

```json
{
  "mode": "search",
  "query": "CRISPR AND 2024[PDAT]",
  "pmids": [],
  "retmax": 500,
  "retstart": 0,
  "sort": "relevance",
  "dateFrom": "",
  "dateTo": "",
  "publicationTypes": [],
  "includeAbstract": true,
  "includeMeSH": true,
  "email": "apify-pubmed-scraper@example.com",
  "tool": "apify-pubmed-scraper",
  "apiKey": "",
  "maxItems": 500,
  "requestDelayMs": 400
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "pmids": [],
    "publicationTypes": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("aurenic/pubmed-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "pmids": [],
    "publicationTypes": [],
}

# Run the Actor and wait for it to finish
run = client.actor("aurenic/pubmed-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "pmids": [],
  "publicationTypes": []
}' |
apify call aurenic/pubmed-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,aurenic/pubmed-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/45PRIPGnAdgekL9dE/builds/JovxeRTkKK1Lk26mh/openapi.json
