# Research Papers MCP: PubMed, arXiv, Crossref (`digital_influx/research-papers-mcp`) Actor

MCP server for Claude, ChatGPT, Gemini and Cursor: search 180M+ research papers (Crossref, PubMed and preprints via Europe PMC, arXiv), read abstracts, references and citing papers, and open access full texts.

- **URL**: https://apify.com/digital_influx/research-papers-mcp.md
- **Developed by:** [Bruno Petrelli](https://apify.com/digital_influx) (community)
- **Categories:** MCP servers, AI, Education
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 paper listeds

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Research Papers MCP: PubMed, arXiv, Crossref

**An MCP server that gives Claude, ChatGPT, Gemini, Cursor and any other MCP client seven research tools over 180M+ scholarly works:**

| Tool | What the agent gets |
|---|---|
| `search_papers` | Search Crossref, the DOI registry of 180+ million journal articles, conference papers, books, preprints and reports from every field, by words, author, journal, years and type, by relevance, newest or most cited. |
| `search_biomedical` | Search the life-science literature of Europe PMC: every PubMed record, PubMed Central and preprints (bioRxiv, medRxiv, Research Square...), 45+ million records, with abstracts, citation counts and open access status. |
| `search_arxiv` | Search arXiv preprints (AI and computer science, physics, mathematics, statistics, quantitative biology, economics) by words, category, author and years. |
| `get_paper` | Everything about one paper: all authors with ORCID and affiliations, the whole abstract, keywords and MeSH terms, citation and reference counts, funders, license and links. |
| `get_references` | The reference list of a paper, in its order, with the DOI or PubMed ID of each cited work. |
| `get_citations` | The papers that cite a paper, newest or most cited first. |
| `read_full_text` | An open access article in full, as text with its sections and tables, when its license allows reuse. |

Ask your agent things like *"What has been published this year on GLP-1 receptor agonists and dementia? Summarize the evidence"*, *"Find the most cited papers on microplastics in drinking water"*, *"What are the newest arXiv papers on LLM agents?"*, *"Who followed up on this CRISPR study, and did anyone contradict it?"* or *"Read the methods of this trial and list its limitations"*: it picks the tool and gets clean JSON back. Every result carries an `id` (DOI, PubMed ID or arXiv ID) that the other tools take. **You pay only for results:** a search that finds nothing, an unknown id or an article that cannot be read costs nothing.

### Connect your agent

**Easiest, with your Apify account (sign-in in the browser, no token to copy):** add this remote MCP server URL to your client:

```
https://mcp.apify.com?tools=digital_influx/research-papers-mcp
```

- **Claude** (claude.ai, Claude Desktop): add a custom connector with that URL.
- **ChatGPT:** add it as a connector (MCP server URL) in developer mode.
- **Cursor, VS Code, Windsurf, Gemini CLI and other clients:** add a remote MCP server. With a token instead of the browser sign-in:

```json
{
  "mcpServers": {
    "research-papers": {
      "url": "https://mcp.apify.com?tools=digital_influx/research-papers-mcp",
      "headers": { "Authorization": "Bearer YOUR_APIFY_TOKEN" }
    }
  }
}
```

(Gemini CLI writes `httpUrl` instead of `url` in `~/.gemini/settings.json`.)

**Directly, without the Apify MCP server:** this Actor runs in Standby mode as a Streamable HTTP MCP server. Take its Standby URL from the Actor's API tab in Apify Console, add `/mcp`, and send your Apify token as `Authorization: Bearer YOUR_APIFY_TOKEN`.

**A record of every call:** each tool call your agent makes is also saved as a row of the Standby run's dataset (tool, arguments, a one-line summary and the result), so you can check later what it asked and got. **Without an agent:** a standard run calls one tool from the input (`tool` and `arguments`) and saves the result to the dataset, so the tools also work from the API, schedules and integrations (Make, n8n, Zapier).

### The tools

#### search_papers

```json
{ "query": "microplastics", "fromYear": 2020, "toYear": 2021, "type": "journal-article", "sort": "cited", "limit": 10 }
```

All arguments are optional, but give at least a `query`, an `author` or a `journal`: `fromYear`, `toYear`, `type` (`journal-article`, `proceedings-article`, `posted-content` for preprints, `book-chapter`, `book`, `report`, `dissertation`, `dataset`), `onlyWithAbstract`, `sort` (`relevance`, `newest`, `cited`) and `limit` (10 by default, at most 50). Returns title, first authors, year, journal, citation count, DOI and the abstract when the publisher deposited it (first 1,000 characters). Crossref matches words, not meaning: add an author, a journal or years, or sort by citations, to see the landmark papers first. Crossref also matches any word of an author's name, so with `author` the tool reads more results and keeps only the works with an author of that surname (and initial, when you give the first name): `"Geoffrey Hinton"` sorted by citations gives *Deep learning*, *ImageNet classification with deep convolutional neural networks* and *Learning representations by back-propagating errors*, without the papers of other Geoffreys.

Real result (2026-10-04): the most cited microplastics papers of 2020-2021 are *Plasticenta: First evidence of microplastics in human placenta* (Environment International, 3,218 citations) and *Environmental exposure to microplastics: An overview on possible human health effects* (2,495).

#### search_biomedical

```json
{ "query": "GLP-1 receptor agonists AND dementia", "openAccessOnly": false, "sort": "relevance", "limit": 10 }
```

The query takes Europe PMC syntax: phrases in quotes, `AND` / `OR` / `NOT`, and fields such as `TITLE:`, `ABSTRACT:` or `MESH:`. Optional: `author` (`"Doudna JA"` or a surname), `journal`, `fromYear`, `toYear`, `openAccessOnly`, `preprintsOnly`, `reviewsOnly`, `sort` and `limit`. Each paper comes with its PubMed ID, PMC ID and DOI, journal, year, citation count, publication types, open access status, whether `read_full_text` can read it, and the abstract.

Real result (2026-10-04): 3,109 papers and preprints match `GLP-1 receptor agonists AND dementia`; one of them, a 2026 PLOS One cohort study, is CC BY and can be read in full.

#### search_arxiv

```json
{ "query": "retrieval augmented generation", "category": "cs.CL", "sort": "newest", "limit": 10 }
```

Optional: `query` (words in titles and abstracts, phrases in quotes), `category` (an arXiv code: `cs.LG`, `cs.CL`, `math.AP`, `q-bio.NC`, `hep-th`...), `author`, `fromYear`, `toYear`, `sort` (`relevance` or `newest`) and `limit`. Returns title, authors, submission date, version, categories, the comments field (*"Accepted to NeurIPS"*), the abstract and links to the arXiv page and PDF.

#### get_paper

```json
{ "id": "10.1038/nature14539" }
```

`id` is a DOI, a PubMed ID (`23287718`), a PubMed Central ID (`PMC3795411`), a Europe PMC preprint ID (`PPR1314996`) or an arXiv ID (`1706.03762`), bare or as a link. Crossref and Europe PMC are merged: title, all authors with ORCID and affiliations, journal, publisher, volume, issue and pages, the whole abstract, keywords, MeSH terms, citation count (Crossref and Europe PMC), reference count, funders and grants, license, open access status and links. arXiv papers come from DataCite, with the citation count of OpenCitations.

Real results (2026-10-04): *Deep learning* (LeCun, Bengio and Hinton, Nature 2015) has 77,915 citations in Crossref and 103 references; *Attention Is All You Need* is at version 7 (last revised 2023-08-02) with 24,956 citations counted by OpenCitations.

#### get_references

```json
{ "id": "10.1038/nature14539", "limit": 25 }
```

The reference list in the paper's order: title, authors, year, journal, volume, pages, and the DOI or PubMed ID of each cited work, or the citation text when the publisher deposited only that (DOI-only references get their title from Crossref). From Crossref, where most publishers deposit open references, or Europe PMC. `limit`: 25 by default, at most 200.

#### get_citations

```json
{ "id": "23287718", "sort": "cited", "limit": 20 }
```

The papers that cite a paper, from Europe PMC's citation network (life sciences: PubMed, PMC and preprints), `newest` (default) or `cited` first, with optional `fromYear`, `toYear` and `openAccessOnly`. For a paper outside Europe PMC it returns the citation count from Crossref or OpenCitations, free.

Real result (2026-10-04): 11,830 papers cite Cong et al. 2013, *Multiplex genome engineering using CRISPR/Cas systems*; the most cited of them is *Genome engineering using the CRISPR-Cas9 system* (Nature Protocols, 9,863 citations).

#### read_full_text

```json
{ "id": "PMC13502699", "sections": ["methods", "results"], "maxCharacters": 30000 }
```

The full text of an open access article from Europe PMC, as plain text with its headings (`## Methods`, `### Study design`), figure and table captions, and the tables row by row; plus the list of sections with their length, so the agent can ask for only the ones it needs. `maxCharacters`: 30,000 by default, at most 100,000. Only articles whose license allows reuse are returned (CC BY, CC BY-SA, CC0, public domain); for CC BY-NC, CC BY-ND and closed articles the tool gives the link to read it, free.

### How much it costs

Pay per event, only for results:

| Event | When | Price |
|---|---|---|
| `paper` | each paper a search returns, and each reference or citing paper listed | USD 0.002 |
| `paper-details` | each paper read with `get_paper` | USD 0.01 |
| `full-text` | each article read in full with `read_full_text` | USD 0.02 |

A search of 10 papers costs USD 0.02. Searches that find nothing, unknown ids, papers without an open reference list and articles that cannot be read in full cost nothing. Runs in Standby use 256 MB.

### Good to know

- **Respectful by design:** an honest User-Agent (`research-papers-mcp`), only public APIs at the pace each one asks for (Crossref's public pool, one request per second; Europe PMC, 10 seconds between requests as its robots.txt asks, so a call that follows another one there can wait a few seconds), no proxies, no login. arXiv is read through DataCite, where arXiv registers every paper, because arXiv's own API is closed to robots.
- **Licenses are respected.** Bibliographic metadata are facts free to reuse; abstracts remain under the copyright of their authors and publishers and come as Crossref and Europe PMC distribute them. Full texts are returned only under licenses that allow reuse, with title, authors, journal, DOI and license, so the credit goes with the text.
- **Sources:** Crossref, Europe PMC (EMBL-EBI), DataCite and OpenCitations. Citation counts differ between sources because each counts the citing works it knows.
- **More tools for agents** by the same author: Company Research MCP, Job Search MCP, Marketing & SEO Research MCP and US & UK Public Records MCP.

# Actor input Schema

## `tool` (type: `string`):

The tool to call in a standard run. search_papers: 180M+ papers from every field (Crossref). search_biomedical: PubMed, PMC and preprints (Europe PMC). search_arxiv: arXiv preprints. get_paper: full record of a paper. get_references: what a paper cites. get_citations: papers that cite it. read_full_text: an open access article in full.

## `arguments` (type: `object`):

The tool's arguments as JSON. search_biomedical: {"query": "GLP-1 receptor agonists AND dementia", "limit": 5}. search_papers: {"query": "microplastics", "sort": "cited", "limit": 10}. search_arxiv: {"query": "retrieval augmented generation", "category": "cs.CL"}. get_paper, get_references, get_citations and read_full_text: {"id": "10.1038/nature14539"} (a DOI, PubMed ID, PMC ID or arXiv ID).

## Actor input object example

```json
{
  "tool": "search_biomedical",
  "arguments": {
    "query": "GLP-1 receptor agonists AND dementia",
    "limit": 5
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "tool": "search_biomedical",
    "arguments": {
        "query": "GLP-1 receptor agonists AND dementia",
        "limit": 5
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("digital_influx/research-papers-mcp").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "tool": "search_biomedical",
    "arguments": {
        "query": "GLP-1 receptor agonists AND dementia",
        "limit": 5,
    },
}

# Run the Actor and wait for it to finish
run = client.actor("digital_influx/research-papers-mcp").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "tool": "search_biomedical",
  "arguments": {
    "query": "GLP-1 receptor agonists AND dementia",
    "limit": 5
  }
}' |
apify call digital_influx/research-papers-mcp --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,digital_influx/research-papers-mcp"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bbHdqur3JsDgHBxLY/builds/6tHPkP1SX8vkfsLXy/openapi.json
