# Crossref Works Search: Papers, DOIs, Citations, No Login (`conserving_celerytop/crossref-works-search`) Actor

$0.75 per 1,000 works. Search 180+ million scholarly works on Crossref by keywords, author, journal, ISSN, DOI list, date, type and funder. One row per work: DOI, title, authors, journal, publisher, dates, citation count, references, license and links. Abstracts optional. No login or API key.

- **URL**: https://apify.com/conserving\_celerytop/crossref-works-search.md
- **Developed by:** [Don Mangu](https://apify.com/conserving_celerytop) (community)
- **Categories:** Education, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.75 / 1,000 works

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Crossref Works Search

### What the Crossref works search does

This Actor searches more than 180 million scholarly works registered with **Crossref**: journal articles, book chapters, conference papers, preprints, datasets, reports and more. Search by keywords, author, journal name, ISSN, publication date, work type and funder, or paste a list of DOIs to look them up in bulk. You get one row per work with the DOI, title, author names, journal, publisher, publication dates, volume, issue and pages, citation count, number of references, license, funders and full-text links. Abstracts are optional: turn on **Include abstracts** to add the abstract text when the publisher deposited one.

Use it for literature reviews, citation counts for a list of papers, a journal's output for a year, everything a funder paid for, or a clean DOI metadata export for your reference manager, spreadsheet or AI pipeline. The data comes from Crossref's public metadata, read live when you run it. No login, no API key and no Crossref account are needed.

### How to search Crossref works, step by step

1. Open the Actor and type **Search words**, for example `large language models`. Or leave it empty and use **Author**, **Journal name**, **Journal ISSNs** or **Funders** instead.
2. Narrow the results with **Published from** and **Published until** (2024, 2024-05 or 2024-05-17) and **Work types** (journal article, book chapter, conference paper, preprint and others).
3. Choose **Sort by**: relevance, newest, oldest, most cited or recently updated.
4. Set **Maximum results** (1 to 10,000) to cap the run and the cost.
5. To look up known papers instead, paste them into **DOIs** (plain DOIs or doi.org links, up to 10,000). The search fields are not used when DOIs are given.
6. If you need abstracts, open **Advanced options** and turn on **Include abstracts**. Abstracts are off by default.
7. Click **Start**. Results appear in the Overview table and can be downloaded as JSON, CSV, Excel or HTML, or read through the Apify API.

### How much does a Crossref works search cost?

Pay per event: you pay **$0.75 per 1,000 works** ($0.00075 per work) on every plan plus the standard Actor start event. Only rows with work data are charged. DOIs that Crossref does not know (status `not_found`) and entries that are not DOIs (status `invalid_input`) return a free row.

Worked example: a literature scan of 2,000 journal articles costs 2,000 x $0.00075 = $1.50. Citation counts for a reading list of 150 DOIs cost about $0.11. A weekly run of the 500 newest papers on a topic costs about $0.38 per run. Set **Published from** to the date of the previous run, since a work returned in an earlier run is returned and charged again. Set a spending limit on the run and the Actor stops cleanly when the limit is reached.

### Input example for a Crossref works search

```json
{
    "query": "large language models",
    "publishedFrom": "2025-01",
    "types": ["journal-article"],
    "sortBy": "most-cited",
    "maxResults": 500
}
```

DOI lookup:

```json
{
    "dois": ["10.1038/nature12373", "https://doi.org/10.1126/science.1225829"]
}
```

| Field | Console name | What it does |
|---|---|---|
| `query` | Search words | Words to find in titles, abstracts, authors, journals and other metadata |
| `author` | Author | Author name to match |
| `journal` | Journal name | Journal, book series or proceedings title to match |
| `issns` | Journal ISSNs | Works from these journals only |
| `publishedFrom` | Published from | Earliest publication date |
| `publishedUntil` | Published until | Latest publication date |
| `types` | Work types | journal-article, book-chapter, proceedings-article, posted-content and more |
| `funders` | Funders | Funder Registry IDs or exact funder names |
| `dois` | DOIs | Look up these works directly |
| `sortBy` | Sort by | relevance, newest, oldest, most-cited, recently-updated |
| `maxResults` | Maximum results | 1 to 10,000 works from a search |
| `onlyWithAbstract` | Works with an abstract only | Skip works without an abstract |
| `includeAbstracts` | Include abstracts | Add the abstract text to each row (off by default) |

### Output example: one Crossref work

```json
{
    "doi": "10.1038/nature12373",
    "doiUrl": "https://doi.org/10.1038/nature12373",
    "title": "Nanometre-scale thermometry in a living cell",
    "authors": ["G. Kucsko", "P. C. Maurer", "N. Y. Yao", "M. Kubo", "H. J. Noh", "P. K. Lo", "H. Park", "M. D. Lukin"],
    "authorCount": 8,
    "journal": "Nature",
    "issn": ["0028-0836", "1476-4687"],
    "publisher": "Springer Science and Business Media LLC",
    "type": "journal-article",
    "publishedDate": "2013-07-31",
    "publishedYear": 2013,
    "volume": "500",
    "issue": "7460",
    "pages": "54-58",
    "citationCount": 1827,
    "referencesCount": 30,
    "license": "http://www.springer.com/tdm",
    "hasAbstract": false,
    "abstract": null,
    "funders": [],
    "publisherUrl": "https://www.nature.com/articles/nature12373",
    "status": "ok"
}
```

Other fields in every row: `subtitle`, `journalAbbreviation`, `isbn`, `publishedPrintDate`, `publishedOnlineDate`, `licenses`, `subjects`, `links`, `relevanceScore`, `createdDate`, `indexedAt` and `error`. A STATS record in the key-value store gives the number of matching works, the works delivered and the charges for the run.

### Data fields in the Crossref works dataset

- **doi**, **doiUrl**: the DOI in lower case and its doi.org link.
- **title**, **authors**, **authorCount**: the work's title and author names as published. Only names are returned.
- **journal**, **issn**, **publisher**, **type**: where and by whom the work was published.
- **publishedDate**, **publishedYear**: the earliest publication date Crossref holds, print or online.
- **citationCount**: how many works registered with Crossref cite this one. Other citation indexes count other sources, so their numbers differ.
- **referencesCount**: how many references the work cites.
- **license**: the license URL for the version of record, or the first license listed. Some publishers list a text-and-data-mining license here, not an open license.
- **hasAbstract**, **abstract**: whether the publisher deposited an abstract (about a quarter of works), and the abstract as plain text. The text is filled only when you turn on **Include abstracts**; otherwise it is null.
- **funders**: funder names with their Funder Registry ID and award numbers.

### FAQ

#### Is it legal to use Crossref metadata?

Crossref states that almost none of its metadata is subject to copyright and that it may be used for any purpose. Abstracts are the exception: they can remain the publisher's copyright. That is why abstracts are off by default. Turn on **Include abstracts** only if you need them. This Actor reads only public metadata, needs no login and returns author names only, with no contact details.

#### Can I reuse the abstracts?

Abstracts are the publisher's text, and you are responsible for how you reuse them.

#### Which works are included?

Works with a Crossref DOI: most journals, many books and conference proceedings, and preprint servers that register DOIs. Works registered only with another DOI agency are not included. The Actor reads Crossref's open metadata, so there are no CAPTCHAs or blocks.

#### How many results can I get?

Up to 10,000 works per run from a search, and up to 10,000 DOIs per run in a DOI lookup. For a larger set, split it by year or work type and run once per part.

#### Why do my search results include loosely related works?

Search words match any of the words, ranked by relevance. Sort by relevance and use a small **Maximum results**, or add **Work types**, dates or a journal to narrow it down.

#### Why does a DOI return not\_found?

Crossref has no record for it. The DOI may be registered with another DOI agency or may contain a typing error. These rows are free.

#### Why is a funder name rejected?

Funder names must match the Funder Registry name exactly. The error message lists the closest matches with their IDs; paste the ID into **Funders**.

#### Is this Actor affiliated with Crossref?

No. It is an independent tool that reads Crossref's public metadata. Crossref is a trademark of its owner.

### Related tools

Pair it with a reference manager import (download the dataset as CSV or JSON), or schedule the Actor weekly with **Sort by** set to newest and **Published from** set to last week's date to follow new papers on a topic.

# Actor input Schema

## `query` (type: `string`):

Enter words to find in titles, abstracts, authors, journals and other metadata, for example "large language models". Leave empty to search by author, journal, ISSN or funder only.

## `author` (type: `string`):

Enter an author name to match, for example "Jennifer Doudna".

## `journal` (type: `string`):

Enter a journal, book series or proceedings title to match, for example "Nature Communications". For an exact journal use Journal ISSNs.

## `issns` (type: `array`):

Add ISSNs to keep works from those journals only, for example 1476-4687. Several ISSNs are combined with OR.

## `publishedFrom` (type: `string`):

Enter the earliest publication date as 2024, 2024-05 or 2024-05-17.

## `publishedUntil` (type: `string`):

Enter the latest publication date as 2025, 2025-12 or 2025-12-31.

## `types` (type: `array`):

Choose the work types to return. Several types are combined with OR. Leave empty for all types.

## `funders` (type: `array`):

Add funders by Funder Registry ID (100000001 or 10.13039/100000001) or by exact name (National Science Foundation). Several funders are combined with OR.

## `dois` (type: `array`):

Add DOIs or doi.org links to look up those works directly, up to 10,000. When DOIs are given, the search fields and filters are not used. Unknown DOIs return a free row with status not\_found.

## `sortBy` (type: `string`):

Choose the order of the results.

## `maxResults` (type: `integer`):

Set how many works to return from a search, from 1 to 10,000.

## `onlyWithAbstract` (type: `boolean`):

Return only works whose record includes an abstract.

## `includeAbstracts` (type: `boolean`):

Turn on to add the abstract text to each row when the publisher deposited one. Off by default: abstracts are the publisher's text and can be under copyright, so you are responsible for how you reuse them.

## Actor input object example

```json
{
  "query": "large language models",
  "publishedFrom": "2025-01",
  "types": [
    "journal-article"
  ],
  "sortBy": "relevance",
  "maxResults": 20,
  "onlyWithAbstract": false,
  "includeAbstracts": false
}
```

# Actor output Schema

## `works` (type: `string`):

Dataset items, one per work: doi, title, authors, journal, publisher, publishedDate, type, citationCount, referencesCount, license, abstract (when Include abstracts is on), funders, links and status.

## `stats` (type: `string`):

JSON record with works delivered, works matching, requests, retries, failed requests, charges, warnings and errors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "large language models",
    "publishedFrom": "2025-01",
    "types": [
        "journal-article"
    ],
    "maxResults": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("conserving_celerytop/crossref-works-search").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "query": "large language models",
    "publishedFrom": "2025-01",
    "types": ["journal-article"],
    "maxResults": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("conserving_celerytop/crossref-works-search").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "large language models",
  "publishedFrom": "2025-01",
  "types": [
    "journal-article"
  ],
  "maxResults": 20
}' |
apify call conserving_celerytop/crossref-works-search --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,conserving_celerytop/crossref-works-search"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/fMPRYJeJyWd7D5BQ9/builds/oj0CGkQ3sinTb0gw4/openapi.json
