# PubMed Scraper: Articles, Abstracts & Citations (`punkrecordsdata/pubmed-articles-scraper`) Actor

Search PubMed and export biomedical articles with abstracts, MeSH terms, DOIs, journals, cited-by and related papers. Export to CSV, Excel, JSON or XML.

- **URL**: https://apify.com/punkrecordsdata/pubmed-articles-scraper.md
- **Developed by:** [PunkRecordsData](https://apify.com/punkrecordsdata) (community)
- **Categories:** AI, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 article records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

<p align="center">
  <img src="https://api.apify.com/v2/key-value-stores/AAm3a1h3Z9nYfrvh9/records/banner?v=2" alt="PunkRecordsData" width="100%" />
</p>

## 🧫 PubMed Scraper - Articles, Abstracts & Citations - PunkRecordsData

> 🚀 **Export PubMed articles in seconds.** Search 37M+ biomedical papers with full PubMed syntax and get structured rows: title, journal, authors, DOI, PMCID, publication types, full abstract, MeSH terms and keywords, plus cited-by and related articles as their own rows. A "crispr gene editing" search matches 25,500 articles ready to export as CSV, Excel, JSON or XML.

The PubMed Scraper uses NCBI's official E-utilities, the same interface behind the PubMed website, so search behavior matches exactly what researchers expect, including field tags like `smith j[au]` and boolean queries. Four billable modules let you pull exactly the depth you need: bibliographic rows, abstracts with MeSH indexing, citing papers and PubMed's own related-article suggestions.

| 🎯 Target Audience | 💡 Primary Use Cases |
|---|---|
| Medical affairs and literature-review teams | Systematic searches exported straight to screening sheets |
| Pharma competitive intelligence | Track publications around a molecule or indication |
| Bibliometrics and research analysts | Citation snowballing with cited-by rows |
| RAG / medical AI builders | Clean abstracts + MeSH terms for indexing pipelines |

### 📋 What the PubMed Scraper does

- **Article records**: PMID, title, journal, up to 15 authors, publication and e-pub dates, volume/issue/pages, DOI, PMCID, publication types and language.
- **Abstract module**: the full structured abstract, MeSH headings and author keywords per article.
- **Cited-by module**: articles citing each result, one full bibliographic row each.
- **Related module**: PubMed's similar-articles list as rows.
- **Real PubMed search**: full query syntax, best-match or newest-first sorting, and publication-date range filters, plus direct PMID lookups.

> 💡 **Why it matters:** literature reviews die in copy-paste. One run turns a PubMed query into a screening-ready spreadsheet with abstracts and MeSH terms attached, and citation snowballing becomes a checkbox.

### 📊 Output of the PubMed search

Real sample from a live run:

```json
{
  "recordType": "article",
  "pmid": "31295471",
  "title": "CRISPR-Cas9 system: A new-fangled dawn in gene editing.",
  "url": "https://pubmed.ncbi.nlm.nih.gov/31295471/",
  "journal": "Life sciences",
  "authors": ["Bhattacharya S", "..."],
  "pubDate": "2019 Sep 15",
  "doi": "10.1016/j.lfs.2019.116636",
  "abstract": "Genome editing by engineered nucleases has revolutionized biomedical research...",
  "meshTerms": ["Animals", "CRISPR-Cas Systems", "Gene Editing"],
  "error": null
}
```

### ✨ Why choose this PubMed scraper

- **4 billable events** (articles, abstracts, cited-by, related) against the alternatives' 1.
- **MeSH terms and keywords included** with every abstract, the indexing layer most exports drop.
- **Citation snowballing built in**: cited-by rows carry full bibliographic data, not bare IDs.
- **True PubMed syntax**, so saved searches port over unchanged.
- **Honest billing**: the abstract event is charged only when an abstract actually exists.

### 📈 How this PubMed scraper compares to alternatives

Measured against the PubMed actors on the Apify Store (September 2026):

| | This actor | Closest priced alternative | Most-used alternative |
|---|---|---|---|
| Billable data events | 4 | 1 | 1 |
| MeSH terms + keywords | Yes | No | No |
| Cited-by / related rows | Yes | No | No |
| Full PubMed query syntax | Yes | Yes | Partial |
| Price per 1,000 articles | $2.50 | $2.00 | $2.99 |

### 🚀 How to use the PubMed Scraper

1. Create a free Apify account (with $5 of credit) at console.apify.com.
2. Open this actor's page and click **Try for free**.
3. Paste your PubMed query (or PMIDs), set sorting and date range.
4. Toggle abstracts, cited-by and related modules as needed.
5. Click **Start** and download CSV, Excel, JSON or XML.

### 💼 Business use cases

#### Systematic literature reviews

Export the full search with abstracts into your screening tool; inclusion/exclusion starts immediately.

#### Molecule and competitor monitoring

Weekly scheduled runs on drug names, newest first, into a Slack digest.

#### KOL identification

Author fields across hundreds of papers on a topic reveal who publishes most where.

#### Medical AI corpora

Abstracts plus MeSH terms make clean, licensed-source training and retrieval data.

### 🔌 Automating the PubMed Scraper

Connect to **Make**, **Zapier**, **Slack**, **Airbyte**, **GitHub** or **Google Drive** through Apify integrations: scheduled alerts for new papers, warehouse syncs, or webhook-driven review pipelines.

### 🌟 Beyond business use cases

- **Research:** reproducible bibliometric datasets with citation edges.
- **Personal:** follow the literature on a condition that matters to you.
- **Non-profit:** evidence monitoring for guidelines and advocacy.
- **Experimentation:** a structured playground over the official E-utilities.

### 🤖 Ask an AI assistant about this scraper

> "I need PubMed search results with abstracts and MeSH terms as CSV, plus the papers citing each result. Would the PubMed Scraper on Apify (apify.com/punkrecordsdata/pubmed-articles-scraper) handle a weekly schedule?"

### ❓ Frequently Asked Questions

#### 🧫 How do I export PubMed search results to CSV or Excel?

Paste your query, click Start, and download the dataset from the Storage tab in CSV, Excel, JSON or XML.

#### 📄 Does it include full abstracts?

Yes, the abstract module fetches the complete structured abstract plus MeSH headings and keywords, billed only when an abstract exists.

#### 🔎 Can I use PubMed's advanced query syntax?

Yes, everything PubMed accepts works: boolean operators, field tags like \[au], \[ti], \[mh], and phrase quoting.

#### 🕸 How do I get the papers that cite an article?

Enable the cited-by module (or paste the PMID directly); each citing paper becomes a full bibliographic row.

#### 🧬 What are MeSH terms and why do they matter?

Medical Subject Headings, NLM's controlled vocabulary. They make filtering and deduplication far more reliable than keywords alone.

#### 🆔 Can I fetch specific articles by PMID?

Yes, the PMIDs field accepts a list and runs before any searches.

#### 📅 Can I limit by publication date?

Yes, after/before date filters map to PubMed's own date range on publication date.

#### 💵 Do I pay for abstracts that don't exist?

No. Articles without abstracts show "Not Available" and the abstract event is not charged.

#### 📦 How many articles can one run return?

Up to 1,000,000 per run on paid plans (PubMed caps any single query's pagination at 9,999 records; split by date ranges for more). Free users get a 10-article preview.

#### ⚙️ Does it need an NCBI API key?

No. It paces itself under NCBI's no-key limits and identifies itself politely.

#### 🌐 Does it cover full-text PDFs?

It exports the PMCID when a free full-text version exists in PubMed Central, so your pipeline can fetch it; PDFs themselves are out of scope.

### 🔌 Integrate with any app

Datasets are available via the Apify API in JSON, CSV, Excel or XML, with webhooks on completion, ready for Python, R, Sheets or reference managers.

### 🔗 Recommended Actors

- [Clinical Trials Scraper](https://apify.com/punkrecordsdata/clinical-trials-scraper) - the trials behind the papers, with sites and adverse events
- [FDA Drug Safety Scraper](https://apify.com/punkrecordsdata/fda-drug-safety-scraper) - recalls, FAERS and labels for the same molecules
- [arXiv Research Papers Scraper](https://apify.com/punkrecordsdata/arxiv-research-papers-scraper) - preprints with citation metrics
- [NPPES NPI Registry Scraper](https://apify.com/punkrecordsdata/nppes-npi-registry-scraper) - US clinician lookups at scale

> 💡 **Pro Tip:** browse the complete [PunkRecordsData collection](https://apify.com/punkrecordsdata) for more data tools.

**🆘 Need Help?** contact.punkrecordsdata@gmail.com

> **⚠️ Disclaimer:** independent tool, not affiliated with NCBI, NLM or the NIH; only publicly available data. Not medical advice.

# Actor input Schema

## `searchTerms` (type: `array`):

One PubMed search per term. Full PubMed syntax works (e.g. "crispr", "breast cancer AND immunotherapy", "smith j\[au]").

## `pmids` (type: `array`):

Fetch exact articles by PubMed ID. Runs in addition to searches.

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## `sortBy` (type: `string`):

Order of PubMed search results.

## `publishedAfter` (type: `string`):

Only articles published on or after this date.

## `publishedBefore` (type: `string`):

Only articles published on or before this date.

## `includeAbstract` (type: `boolean`):

Fetch the full abstract, MeSH headings and author keywords per article. Billed as its own event.

## `includeCitedBy` (type: `boolean`):

Articles that cite each result (one row per citing paper).

## `maxCitedByPerArticle` (type: `integer`):

Cap on citing-paper rows per article.

## `includeRelated` (type: `boolean`):

PubMed's similar-articles list (one row per related paper).

## `maxRelatedPerArticle` (type: `integer`):

Cap on related-article rows per article.

## Actor input object example

```json
{
  "searchTerms": [
    "crispr gene editing"
  ],
  "pmids": [],
  "maxItems": 10,
  "sortBy": "relevance",
  "includeAbstract": true,
  "includeCitedBy": false,
  "maxCitedByPerArticle": 20,
  "includeRelated": false,
  "maxRelatedPerArticle": 10
}
```

# Actor output Schema

## `overview` (type: `string`):

Key fields per article

## `fullData` (type: `string`):

Complete dataset with all fields

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "crispr gene editing"
    ],
    "pmids": [],
    "maxItems": 10,
    "publishedAfter": "",
    "publishedBefore": ""
};

// Run the Actor and wait for it to finish
const run = await client.actor("punkrecordsdata/pubmed-articles-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerms": ["crispr gene editing"],
    "pmids": [],
    "maxItems": 10,
    "publishedAfter": "",
    "publishedBefore": "",
}

# Run the Actor and wait for it to finish
run = client.actor("punkrecordsdata/pubmed-articles-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "crispr gene editing"
  ],
  "pmids": [],
  "maxItems": 10,
  "publishedAfter": "",
  "publishedBefore": ""
}' |
apify call punkrecordsdata/pubmed-articles-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,punkrecordsdata/pubmed-articles-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZljKZgpl4ryBTgkDQ/builds/j4EZHgLvF9MgIY5V6/openapi.json
