# PubMed Articles Scraper (`xtracto/pubmed-articles`) Actor

Search PubMed and get structured article metadata: title, authors, journal, publication dates, volume and pages, DOI, PMC ID, publication types and citation counts. Keyless NCBI E-utilities API.

- **URL**: https://apify.com/xtracto/pubmed-articles.md
- **Developed by:** [Farhan Febrian Nauval](https://apify.com/xtracto) (community)
- **Categories:** Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.33 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PubMed Articles Scraper — Search Results with Full Citation Metadata

Search PubMed and get structured article records: title, full author list, journal and ISSNs,
publication dates, volume/issue/pages, DOI and PMC ID, publication types and citation counts.

Keyless NCBI E-utilities API. No login, no browser, **no proxy needed**.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `searchQueries` | array | *required* | Full PubMed syntax, e.g. `crispr AND cancer[Title]` |
| `publishedFrom` / `publishedTo` | string | — | `YYYY/MM/DD` |
| `dateType` | string | `pdat` | Which date the filter matches — see below |
| `sort` | string | relevance | `pub_date`, `Author`, `JournalName` |
| `maxItemsPerQuery` | integer | `200` | Capped by the API at 9,999 |
| `apiKey` | string (secret) | — | Free from NCBI; raises 3 req/s to 10 |
| `contactEmail` | string | — | NCBI asks callers to identify themselves |

### Output

```jsonc
{
  "_input": "crispr AND cancer",
  "_source": "S1-eutils-esummary",
  "_scrapedAt": "2026-09-09T13:44:12Z",

  "pmid": "42711491",
  "url": "https://pubmed.ncbi.nlm.nih.gov/42711491/",
  "title": "Chemogenomic maps reveal a PRDX1-dependent iron-damage axis…",

  "authors": ["O'Loughlin TA", "Arab A", …],
  "authorCount": 14,
  "firstAuthor": "O'Loughlin TA", "lastAuthor": "Bhatt DM",

  "journal": "Nat Chem Biol",
  "journalFull": "Nature chemical biology",
  "issn": "1552-4450", "eIssn": "1552-4469", "nlmUniqueId": "101231976",

  "pubDate": "2026 Sep 8",          // ← what a date filter matches
  "ePubDate": "2026 Sep 8",
  "sortPubDate": "2026/09/08 00:00", // ← PubMed's sort key; see below
  "docDate": null,
  "volume": "", "issue": "", "pages": "",
  "language": ["eng"],
  "publicationTypes": ["Journal Article"],

  "doi": "10.1038/s41589-026-02312-z",
  "doiUrl": "https://doi.org/10.1038/s41589-026-02312-z",
  "pmcId": null,
  "elocationId": "doi: 10.1038/s41589-026-02312-z",
  "articleIds": { "pubmed": "42711491", "doi": "…" },

  "pmcRefCount": 12,
  "recordStatus": "PubMed - as supplied by publisher",
  "publicationStatus": "aheadofprint"
}
```

### Three things worth knowing

**`pubDate` and `sortPubDate` are different dates, and the filter matches the first.** PubMed's sort
key reflects later revisions: a 2010 article updated in 2024 sorts as 2024. Filtering a run to 2010
returned rows whose `pubDate` was 2010 for 37 of 40 — while their `sortPubDate` was mostly 2013 and
2014\. Compare against `pubDate`, or the filter will look broken when it is working exactly as
specified. `dateType` lets you filter on the Entrez (`edat`) or modification (`mdat`) date instead.

**PubMed serves at most 9,999 records per query.** `retstart` cannot exceed 9998, and `retmax` is
silently clamped: ask for 10,000 and you get 9,999, no warning. Narrowing by date range is the way
past it — the actor logs the true match count and says so when you have asked for more than can be
served.

**The error path is not valid JSON.** Overrunning the ceiling returns HTTP 200 with an `ERROR` key,
and the message embeds a raw newline inside a JSON string. Strict `json.loads` refuses it outright
with `Invalid control character`, so the actual explanation is lost and the failure looks like a
transport error. The success path parses fine — meaning a parser can pass every test and still fail
on the first real failure. This actor parses leniently and reports the message.

### Rate limits and courtesy

NCBI documents 3 requests a second without a key and 10 with one, and asks callers to identify
themselves so they can get in touch about unusual usage rather than simply blocking it. The actor
paces itself to the documented rate, sends a `tool` identifier, and passes your `contactEmail` when
you give one. Neither is authentication; both are the terms of use.

### Errors

| `_error` | Meaning |
|---|---|
| `invalid_input` | Empty query |
| `no_results` | The search ran and matched nothing |
| `api_error` | A 200 carrying an `ERROR` field — usually the 9,999 ceiling |
| `unexpected_shape` | A 200 without `esearchresult` or `result` |
| `blocked` | Every TLS profile was refused |
| `network_error` | The ladder never reached the server |

If every query fails, the run itself fails rather than reporting success over an empty dataset.

# Actor input Schema

## `searchQueries` (type: `array`):

One search per entry. Full PubMed syntax works, e.g. crispr AND (cancer\[Title]) or "machine learning"\[Title/Abstract].

## `publishedFrom` (type: `string`):

Earliest publication date, YYYY/MM/DD. Also the practical way past the 9,999-record ceiling.

## `publishedTo` (type: `string`):

Latest publication date, YYYY/MM/DD.

## `dateType` (type: `string`):

pdat matches the publication date (the `pubDate` field in the output). edat is the date the record entered PubMed, mdat the date it was last modified. Note that `sortPubDate` is a different date again — PubMed's sort key, which reflects later revisions — so a pdat filter will not line up with it.

## `sort` (type: `string`):

PubMed sort key. Leave empty for relevance.

## `maxItemsPerQuery` (type: `integer`):

PubMed's ESearch serves at most 9,999 records for any one query however many matched.

## `apiKey` (type: `string`):

Optional and free from NCBI. Raises the documented rate from 3 requests a second to 10, so large runs finish sooner.

## `contactEmail` (type: `string`):

Optional. NCBI asks callers to identify themselves so they can get in touch about unusual usage rather than simply blocking it.

## `proxyConfiguration` (type: `object`):

Optional. This is a keyless public API with no anti-bot layer, so a proxy is not needed.

## Actor input object example

```json
{
  "searchQueries": [
    "crispr AND cancer"
  ],
  "dateType": "pdat",
  "sort": "",
  "maxItemsPerQuery": 200
}
```

# Actor output Schema

## `overview` (type: `string`):

Dataset items shown in the 'Articles' view.

## `items` (type: `string`):

Every record this run produced, with all fields, as JSON.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "crispr AND cancer"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("xtracto/pubmed-articles").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchQueries": ["crispr AND cancer"] }

# Run the Actor and wait for it to finish
run = client.actor("xtracto/pubmed-articles").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "crispr AND cancer"
  ]
}' |
apify call xtracto/pubmed-articles --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,xtracto/pubmed-articles"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/9ZQVaLL1fkL3pOIxs/builds/GTbyqWtNyeY51MjQE/openapi.json
