# PubMed API - Research Paper Scraper via Europe PMC Search (`captainhandsome/europe-pmc-paper-search`) Actor

Search Europe PMC and PubMed for research papers by topic, author, journal and year. A scientific literature API that filters for open access or abstract-bearing papers and exports PMIDs, DOIs, abstracts, citation counts and full-text links as a structured dataset.

- **URL**: https://apify.com/captainhandsome/europe-pmc-paper-search.md
- **Developed by:** [Joseph McRell](https://apify.com/captainhandsome) (community)
- **Categories:** Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.10 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Europe PMC & PubMed Research Paper Search

Search Europe PMC and PubMed literature by topic, author, journal, year and open-access status. Export identifiers, abstracts, citations and full-text links. This Actor turns the official Europe PMC REST API into bounded, structured datasets for analysis, enrichment, monitoring, and AI-agent workflows.

### What data can I extract?

79 flat columns per paper, every one of them straight from the official Europe PMC REST API:

- **Identifiers** - PMID, PMCID, DOI, NIH manuscript id, journal ISSN, e-ISSN and NLM ID
- **Citation** - journal, abbreviation, volume, issue, pages, print and electronic publication dates
- **Authorship** - full author string, first and last author, author count, ORCID iDs and every institutional affiliation
- **Subject indexing** - MeSH terms, MeSH major topics, author keywords and indexed chemical substances
- **Funding** - funding agencies, grant IDs and grant count
- **Impact and access** - citation count, open-access flag, licence, best full-text availability, PDF link and every full-text URL
- **Integrity** - retraction flag, retraction and erratum citations, and all comment/correction links
- **Preprints and patents** - preprint server and version, plus patent country, type, application number, date and IPC/EPO classifications
- **Mined entities (optional)** - genes and proteins, diseases, organisms, chemicals, Gene Ontology terms, experimental methods and accession numbers

Every run writes flat records to the default dataset. Download results as JSON, CSV, Excel, XML, or access them through the Apify API.

#### All 79 columns

**Identity** — `author_orcids`, `doi`, `entity_accession_numbers`, `first_author_full_name`, `first_author_orcid`, `grant_ids`, `id`, `journal_nlm_id`, `manuscript_id`, `patent_application_number`, `pmcid`, `pmid`, `title`, `version_number`

**Status** — `comment_correction_types`, `has_data`, `has_pdf`, `has_references`, `has_supplementary_material`, `has_text_mined_terms`, `is_author_manuscript`, `is_open_access`, `is_retracted`, `patent_classifications`, `patent_type`, `publication_status`, `publication_types`

**Dates** — `date_of_revision`, `electronic_publication_date`, `embargo_date`, `first_index_date`, `issue`, `patent_application_date`, `print_publication_date`, `pub_year`, `publication_date`, `publisher`

**Location** — `abstract`, `patent_country`, `retraction_notice`

**People** — `author_count`, `authors`, `first_author`, `last_author`

**Contact** — `data_link_tags`, `full_text_url`, `full_text_urls`, `pdf_url`, `url`

**Counts and measures** — `cited_by_count`, `entity_count`, `grant_count`

**Other detail** — `affiliation`, `affiliations`, `chemicals`, `database_cross_references`, `eissn`, `entity_chemicals`, `entity_diseases`, `entity_experimental_methods`, `entity_genes`, `entity_go_terms`, `entity_organisms`, `erratum_notice`, `full_text_availability`, `funders`, `in_epmc`, `issn`, `journal`, `journal_abbrev`, `keywords`, `language`, `license`, `mesh_major_topics`, `mesh_terms`, `pages`, `pub_model`, `source`, `volume`

### Input example

```json
{
  "query": "large language models in medicine",
  "year_from": 2022,
  "open_access_only": true,
  "max_items": 100
}
```

Set `include_entities` to `true` to add the mined-entity columns. They come from the Europe PMC annotations API, batched eight articles per request, so the extra cost is roughly one request per eight results rather than one per record.

`max_items` is a hard output and billing ceiling. The default input is intentionally limited to 10 records so Store tests and first runs stay inexpensive.

### Output example

```json
{
  "id": "42658322",
  "source": "MED",
  "pmid": "42658322",
  "pmcid": "PMC13522008",
  "doi": "10.1007/s10544-026-00843-9",
  "title": "Intranasal CRISPR lipid nanoparticles targeting MAPK9 attenuate neuroinflammation after traumatic brain injury.",
  "authors": "Kara G, Holcomb M, Hijazi AA, Ali Y, López-Espinosa J, Cruz-Pineda L, Park P, Flinn H, Taylor N, Galbraith T, McMahon L, Rostomily R, Leonard F, Villapol S.",
  "first_author": "Kara G",
  "journal": "Biomedical microdevices",
  "journal_abbrev": "Biomed Microdevices",
  "issn": "1387-2176",
  "volume": "28",
  "issue": "3",
  "pages": "58",
  "pub_year": "2026",
  "publication_date": "2026-08-27",
  "publication_types": "research-article, Journal Article",
  "keywords": "Macrophages, Traumatic brain injury, Microglia, neuroinflammation, Intranasal Delivery, Lipid Nanoparticles, Crispr-cas12a, Mapk9",
  "abstract": "Traumatic brain injury (TBI) induces a sustained neuroinflammatory response involving activated microglia and infiltrating myeloid cells, contributing to secondary brain damage and long-term neurological dysfunction. Modulating these inflammatory responses toward a more reparative phenotype represen",
  "affiliation": "Department of Neurosurgery and Center for Neuroregeneration, Houston Methodist Research Institute, Houston, TX, USA.",
  "cited_by_count": 3,
  "is_open_access": true,
  "license": "cc by-nc-nd",
  "has_pdf": true,
  "full_text_url": "https://doi.org/10.1007/s10544-026-00843-9",
  "url": "https://europepmc.org/article/MED/42658322",
  "eissn": "1572-8781",
  "journal_nlm_id": "100887374"
}
```

51 further columns are omitted here for length — the full list is above, and every column appears in the export whether or not the source populated it.

### Use with AI agents and MCP

Apify's hosted MCP server can discover and call this Actor. A suitable agent request is:

> Find 100 open-access biomedical papers about large language models in medicine published since 2022.

Use this exact Actor input:

```json
{
  "query": "large language models in medicine",
  "year_from": 2022,
  "open_access_only": true,
  "max_items": 100
}
```

The Actor succeeds with a nonempty default dataset and exposes its default dataset through the top-level Output schema.

### Pricing and cost control

Output is billed per result at **$0.003 per result** (about $3.00 per 1,000 results), plus a $0.0005 Actor-start charge billed once per gigabyte of memory at run start. Use `max_items` to cap both output volume and charges. The price shown on the Apify Store listing is authoritative.

Use `max_items` to cap returned and billable records. Invalid input is rejected before unnecessary work wherever possible.

### Common use cases

- Source-specific research and market intelligence
- Structured exports for spreadsheets, warehouses and BI systems
- Entity enrichment and monitoring pipelines
- Retrieval and data collection by AI agents

### Reliability

The Actor uses bounded pagination, retries transient upstream failures, deduplicates records where the source exposes stable identifiers, and fails explicitly when the source cannot provide usable output. Production default-input canaries verify a nonempty structured dataset in under five minutes.

### Limitations and responsible use

- Coverage, field availability and update timing are controlled by the upstream public source.
- Optional fields can be null or absent when the source does not publish them.
- This Actor does not bypass authentication, access controls, CAPTCHAs, or source rate limits.
- Customers remain responsible for lawful use, applicable source terms, and restrictions on downstream decisions.

### FAQ

#### What is the best way to run this Europe PMC paper search Actor?

Start with the 10-record default, inspect the dataset, then increase `max_items` and narrow the available filters for your use case.

#### Do I need my own API key?

No external API key is required unless the Input tab explicitly says otherwise. Apify credentials are used normally when invoking the Actor through Apify APIs or MCP.

#### Can an AI agent call it?

Yes. The strict input schema acts as the tool signature, and the dataset plus Output schemas describe the returned records.

#### Can I export the results?

Yes. Apify datasets support JSON, CSV, Excel, XML and API retrieval.

### Support and changes

Open an issue on the Actor page with a redacted input and run ID. See [CHANGELOG.md](./CHANGELOG.md) for contract and maintenance updates.

# Actor input Schema

## `query` (type: `string`):

Keywords, organization name, topic, condition, intervention, sponsor or other source-supported search expression.

## `author` (type: `string`):

Author name filter supported by Europe PMC.

## `journal` (type: `string`):

Journal title filter supported by Europe PMC.

## `year_from` (type: `integer`):

Return papers published in this year or later.

## `year_to` (type: `integer`):

Return papers published in this year or earlier.

## `open_access_only` (type: `boolean`):

Return only records marked open access by Europe PMC.

## `has_abstract_only` (type: `boolean`):

Return only papers for which Europe PMC provides an abstract.

## `sort_by` (type: `string`):

Order results by relevance, citation count or publication date when supported.

## `include_entities` (type: `boolean`):

Add named entities mined from the full text by Europe PMC - genes and proteins, diseases, organisms, chemicals, Gene Ontology terms, experimental methods and accession numbers. Uses the Europe PMC annotations API, batched eight articles per request, so it adds roughly one extra request per eight results.

## `max_items` (type: `integer`):

Hard maximum number of dataset records returned and billed.

## Actor input object example

```json
{
  "query": "CRISPR gene editing",
  "open_access_only": false,
  "has_abstract_only": false,
  "sort_by": "relevance",
  "include_entities": false,
  "max_items": 10
}
```

# Actor output Schema

## `results` (type: `string`):

One flat row per paper.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "CRISPR gene editing"
};

// Run the Actor and wait for it to finish
const run = await client.actor("captainhandsome/europe-pmc-paper-search").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "query": "CRISPR gene editing" }

# Run the Actor and wait for it to finish
run = client.actor("captainhandsome/europe-pmc-paper-search").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "CRISPR gene editing"
}' |
apify call captainhandsome/europe-pmc-paper-search --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,captainhandsome/europe-pmc-paper-search"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/rPOVejpAcoLgwwgj8/builds/byPqHKLzWuVaoJAPk/openapi.json
