# Research Paper Search: DOI Metadata and Citations (`pistachio_implementation/research-paper-search`) Actor

Search 150 million scholarly papers, books and preprints by topic, or look up DOIs. Returns title, authors, journal, publisher, date, abstract, citation count, license, funders and full text links from Crossref's open API. A Google Scholar alternative with no blocking.

- **URL**: https://apify.com/pistachio\_implementation/research-paper-search.md
- **Developed by:** [Hay Equipos](https://apify.com/pistachio_implementation) (community)
- **Categories:** AI, Developer tools, News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 paper saveds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Research Paper Search: DOI Metadata and Citations

Search more than 150 million scholarly works (journal articles, conference papers, books, chapters, preprints, reports and datasets) by topic, or look up any DOI. Each row gives you the title, authors, journal, publisher, publication date, abstract, citation count, reference count, subjects, ISSN, volume, issue and pages, license, funders and links to the full text.

The data comes from Crossref, the registry that issues DOIs for most of the world's publishers, through its open public API. There is no page scraping, no captcha and no blocking, so runs are steady. It is a practical alternative to Google Scholar scrapers when you need clean, structured metadata at volume.

### What you can use it for

- **Literature reviews:** pull every journal article on a topic since 2024, with abstracts, ready for screening or an AI summary.
- **Citation analysis:** sort by most cited to find the key papers in a field, or look up the citation count of your own DOIs.
- **Reference cleanup:** turn a list of DOIs into complete, correctly formatted metadata.
- **Research and grant intelligence:** see which funders appear on recent papers in a field, and which journals publish the most on it.
- **RAG and AI agents:** feed abstracts and metadata into a knowledge base, or let an agent answer "find the 20 most cited papers on CRISPR off target effects since 2022".

### Input

| Field | What it does | Default |
|---|---|---|
| Search queries | Topics, one per line. Matched against titles, authors, journals and other bibliographic text | none |
| DOIs to look up | DOIs or DOI links, one per line | none |
| Published from / until | Date range, as `2024`, `2024-06` or `2024-06-30` | none |
| Work types | Journal article, conference paper, book chapter, book, preprint, report, dissertation, dataset and more | all |
| Sort by | Relevance, newest first or most cited first. Newest and most cited reorder the most relevant matches (five times the results you ask for, at least 200, at most 1,000) so results stay on topic | relevance |
| Only works with an abstract | Skip works without a deposited abstract | off |
| Only works with a license | Keep works that carry a license statement | off |
| Include abstracts | Turn off for smaller rows | on |
| Results per query | Up to 10,000 | 100 |
| Your contact email | Optional. Crossref serves requests with a contact email from a faster pool. Sent only to Crossref, never saved | none |
| Maximum papers | Stop after this many rows in total | 1,000 |

Example input:

```json
{
  "queries": ["large language models in education"],
  "dois": ["10.1038/nature14539"],
  "fromDate": "2024",
  "types": ["journal-article"],
  "onlyWithAbstract": true,
  "sortBy": "relevance",
  "maxResultsPerQuery": 200
}
```

### Output

One row per work. Download as JSON, CSV, Excel or HTML, or read it through the API.

```json
{
  "query": "10.1038/nature14539",
  "source": "doi",
  "doi": "10.1038/nature14539",
  "doiUrl": "https://doi.org/10.1038/nature14539",
  "title": "Deep learning",
  "type": "journal-article",
  "journal": "Nature",
  "publisher": "Springer Science and Business Media LLC",
  "publishedDate": "2015-05-27",
  "year": 2015,
  "authors": ["Yann LeCun", "Yoshua Bengio", "Geoffrey Hinton"],
  "authorCount": 3,
  "abstract": null,
  "citationCount": 77191,
  "referenceCount": 103,
  "issn": ["0028-0836", "1476-4687"],
  "volume": "521",
  "issue": "7553",
  "pages": "436-444",
  "language": "en",
  "licenseUrls": ["https://www.springer.com/tdm"],
  "openLicense": false,
  "funders": [],
  "publisherUrl": "https://www.nature.com/articles/nature14539",
  "scrapedAt": "2026-09-27T06:57:07.221Z"
}
```

`citationCount` is the number of other Crossref works that cite this one. `openLicense` is true when the work carries a Creative Commons license. A DOI that does not exist is listed in `RUN_SUMMARY` in the run's key value store and costs nothing. The same paper is never saved twice in one run.

### Pricing

Pay per event, no subscription, no charge for platform usage on top.

| Event | Price |
|---|---|
| Paper saved | $0.001 (one dollar per 1,000 papers) |

Examples: 200 papers on a topic cost $0.20. Looking up 1,000 DOIs costs $1.00. Set a maximum charge per run in Apify and the actor stops cleanly when it is reached.

### Limits

- Abstracts exist only when the publisher deposited one with Crossref; many older and some large publishers do not. Use "Only works with an abstract" if you need them.
- Citation counts come from Crossref's own citation links. They are usually lower than Google Scholar's, which also counts web pages and theses.
- Without a contact email, Crossref's public pool allows about one request per second; each request returns up to 200 papers, so 10,000 papers take a few minutes. With a contact email it is faster.
- Relevance ranking is Crossref's. Results far down a very broad query get looser; narrow the query or add a date range and type filter.
- "Most cited first" and "newest first" rank the most relevant matches (up to 1,000), not every work in Crossref that shares a word with the query. This is deliberate: Crossref's own citation sort ignores relevance and returns off topic papers.
- Full text is not downloaded. The actor returns links; access depends on the publisher and the license.
- Google Scholar is not used or scraped.

### FAQ

**Do I need an API key or an account?** No. Crossref's REST API is public and free. The optional contact email only speeds things up.

**Is this the same as Google Scholar?** No. It covers works registered with a DOI, which is nearly all journal and conference literature, many books and a growing share of preprints. It does not cover random PDFs on the web.

**Can I search by author?** Put the author's name in the query; bibliographic search matches author names too. The actor does not build profiles of people; author names appear only as the paper's published byline.

**How fresh is it?** Live. Crossref indexes new DOIs within hours of registration.

**Is this affiliated with Crossref?** No. It is an independent tool that reads Crossref's public API.

# Actor input Schema

## `queries` (type: `array`):

Topics to search, one per line, for example "large language models in education" or "CRISPR off target effects". Matched against titles, authors, journals and other bibliographic text.

## `dois` (type: `array`):

DOIs or DOI links, one per line, for example 10.1038/nature14539 or https://doi.org/10.1145/3442188.3445922.

## `fromDate` (type: `string`):

Keep works published on or after this date. Formats: 2024, 2024-06 or 2024-06-30.

## `untilDate` (type: `string`):

Keep works published on or before this date. Formats: 2025, 2025-12 or 2025-12-31.

## `types` (type: `array`):

Keep only these kinds of works. Empty means all.

## `sortBy` (type: `string`):

Relevance to the query, newest first, or most cited first. Newest and most cited reorder the most relevant matches (five times the results you ask for, at least 200 and at most 1,000), so the results stay on topic.

## `onlyWithAbstract` (type: `boolean`):

Skip works whose publisher did not deposit an abstract.

## `onlyOpenLicense` (type: `boolean`):

Keep only works that carry a license statement. Check the openLicense column for Creative Commons.

## `includeAbstract` (type: `boolean`):

Turn off for smaller rows.

## `maxResultsPerQuery` (type: `integer`):

How many works to save for each query.

## `contactEmail` (type: `string`):

Crossref serves requests that include a contact email from a faster pool. It is sent only to Crossref, as they ask, and is never saved in the output.

## `maxItems` (type: `integer`):

Stop after saving this many papers in total.

## Actor input object example

```json
{
  "queries": [
    "large language models in education"
  ],
  "sortBy": "relevance",
  "onlyWithAbstract": false,
  "onlyOpenLicense": false,
  "includeAbstract": true,
  "maxResultsPerQuery": 100,
  "maxItems": 1000
}
```

# Actor output Schema

## `results` (type: `string`):

All rows the run saved to the default dataset.

## `summary` (type: `string`):

The RUN\_SUMMARY record: counts and problems for the whole run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "large language models in education"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pistachio_implementation/research-paper-search").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": ["large language models in education"] }

# Run the Actor and wait for it to finish
run = client.actor("pistachio_implementation/research-paper-search").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "large language models in education"
  ]
}' |
apify call pistachio_implementation/research-paper-search --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pistachio_implementation/research-paper-search"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WfGxbLNvpOXtHnshV/builds/t76khWVSyOlgVmvrf/openapi.json
