# Crossref Academic Paper Metadata Scraper (`parseforge/crossref-academic-paper-scraper`) Actor

Scrapes academic paper metadata from Crossref by search query or publication type. Returns DOI, title, authors, journal, citation counts, and dates.

- **URL**: https://apify.com/parseforge/crossref-academic-paper-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Education, Business, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.62 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner.jpg)](https://apify.com/parseforge?fpr=vmoqkp)

### Crossref Academic Paper Metadata Scraper

**Scrape Crossref academic paper metadata for any DOI, search query, or publication type, up to a million records per run.** Each record includes title, authors, journal, DOI, citation counts, and publication dates. No API key required. Export to CSV, JSON, Excel, or XML.

The Actor queries the public Crossref REST API, filters by publication type, search query, and sort order, and returns each matching scholarly work as one flat row. It covers journal articles, books, dissertations, datasets, and more.

| Who uses it | What they scrape Crossref for |
|---|---|
| Academic researchers | Building a bibliography for a literature review |
| Librarians | Verifying citation metadata for institutional repositories |
| Data scientists | Gathering a corpus of paper metadata for analysis |
| Journal editors | Checking DOI registration and metadata completeness |
| PhD students | Collecting references for a thesis |

### What it does

This Actor collects academic paper metadata from Crossref by search query or publication type, and returns each work as a flat row with DOI, title, authors, journal, citation counts, and dates.

- 🔎 **Search query:** free-text search across titles and other metadata fields.
- 📚 **Publication type filter:** journal-article, book-chapter, book, proceedings-article, dissertation, report, standard, dataset, or posted-content.
- 📅 **Sort order:** by publication date or relevance to the search query.
- 📦 **Bulk export:** up to 1,000,000 records per run for paid users.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with Crossref data

**📚 Build a literature review bibliography.**

A researcher enters a search query like "climate change migration" and exports all matching journal articles with full metadata to CSV for reference management.

**🔍 Verify DOI metadata for a journal issue.**

An editor filters by journal-article and searches for a specific title to confirm authors, volume, issue, and page numbers before publication.

**📊 Analyze publication trends.**

A data scientist fetches all journal articles sorted by published date for a given year and uses citation counts to identify influential papers.

**🎓 Collect references for a thesis.**

A PhD student searches for a topic, filters to dissertations and journal articles, and exports the metadata to build a reference list.

### Why choose this scraper

| | What you get |
|---|---|
| **No API key** | Uses the public Crossref REST API without authentication |
| **Rich metadata** | Returns DOI, title, authors, journal, citation counts, and dates |
| **Flexible filtering** | Search by text or filter by scholarly work type |
| **Scalable** | Fetch up to a million records per run |

### How it compares

This Actor focuses on flexible search and filtering of Crossref metadata with no API key required, similar to other Crossref scrapers but with a simpler input schema.

| Feature | ParseForge | CrossRef Academic Metadata Scraper | Crossref Academic Paper Search | Crossref Api Scraper |
|---|---|---|---|---|
| Search by text query | Yes | Yes | Yes | Yes |
| Filter by publication type | Yes | Not listed | Not listed | Yes |
| Sort by relevance | Yes | Not listed | Not listed | Not listed |
| Fetch up to 1,000,000 records | Yes | Not listed | Not listed | Not listed |
| No API key required | Yes | Not listed | Not listed | Yes |

### What a Crossref record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

```json
{
 "DOI": "10.1157/13053466",
 "type": "journal-article",
 "title": "Comentario: Prevención de los factores de riesgo de los trastornos de la conducta alimentaria en adolescentes:",
 "containerTitle": "Atención Primaria",
 "shortContainerTitle": "Aten Primaria",
 "publisher": "Elsevier BV",
 "issue": "7",
 "volume": "32",
 "page": "408-409",
 "publishedPrintDate": "2203-10",
 "issuedDate": "2203-10",
 "indexedDateTime": "2025-05-30T05:22:24Z",
 "indexedTimestamp": 1748582544347,
 "createdDateTime": "2003-11-17T16:44:58Z",
 "createdTimestamp": 1069087498000
}
```

Every value above comes from a real run. A field a record does not have comes back as `null`.

### Configure the run

Drive the Actor with a search query or leave it empty to fetch the most recently published journal articles. Filter by publication type and sort by published date or relevance. The Input tab lists every parameter.

A first run with the defaults:

```json
{
 "maxItems": 10,
 "filterType": "journal-article",
 "sortOrder": "published"
}
```

A larger pull:

```json
{
 "maxItems": 200,
 "filterType": "journal-article",
 "sortOrder": "published"
}
```

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Open the [Crossref Academic Paper Metadata Scraper](https://apify.com/parseforge/crossref-academic-paper-scraper?fpr=vmoqkp).
3. Set your inputs and any filters, then click **Start**.
4. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to Crossref through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/crossref-academic-paper-scraper"
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Check your search query for typos or overly specific terms. Try a broader query or leave the search field empty to fetch recent articles.

**Why are some fields empty?**

Crossref metadata depends on what publishers deposit. Some fields may be missing for certain records.

**Why is my run limited to 10 items?**

Free users are limited to 10 items as a preview. Upgrade to a paid plan to fetch up to 1,000,000 records.

**Why does sorting by relevance not work?**

Relevance sorting requires a search query. If the query is empty, results are sorted by publication date.

### FAQ

| Question | Answer |
|---|---|
| Do I need a Crossref API key? | No, this Actor uses the public Crossref REST API without authentication. You can start scraping immediately. |
| What metadata fields are returned? | Each record includes DOI, type, title, containerTitle, shortContainerTitle, publisher, issue, volume, page, publishedPrintDate, issuedDate, indexedDateTime, indexedTimestamp, createdDateTime, createdTimestamp, depositedDateTime, depositedTimestamp, isReferencedByCount, referencesCount, URL, ISSN, issnPrint, issnElectronic, licenseUrl, licenseContentVersion, licenseDelayInDays, linkUrl, linkContentType, linkContentVersion, linkIntendedApplication, resourcePrimaryUrl, source, member, prefix, score, alternativeId, journalIssue, authors, scrapedAt, error. |
| Can I search by DOI? | Yes, you can enter a DOI as a search query to retrieve metadata for a specific paper. |
| What publication types are supported? | Journal articles, book chapters, books, proceedings articles, dissertations, reports, standards, datasets, and posted content. |
| How many records can I fetch? | Free users are limited to 10 items as a preview. Paid users can fetch up to 1,000,000 records per run. |
| Can I sort results? | Yes, sort by publication date (newest first) or by relevance to your search query. |
| Is the data from Crossref complete? | Crossref contains metadata for over 150 million scholarly works, but some records may have missing fields depending on what publishers deposited. |
| Can I export to Excel? | Yes, you can export results to CSV, JSON, Excel, or XML. |
| Does this Actor get abstracts? | The dataset does not include an abstracts field. |
| How do I cite the data? | You should cite the original papers, not the metadata. Crossref metadata is provided under a CC0 license. |

### Related actors

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Crossref. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

### 💰 How much does it cost to scrape Crossref Academic Paper Metadata?

This Actor uses **pay-per-result** pricing: **$0.004 per result** collected. You are billed only for the results you receive, so a run that returns nothing costs nothing.

# Actor input Schema

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## `searchQuery` (type: `string`):

A text query to search for in paper titles, abstracts, and other metadata fields. Leave empty to fetch the most recently published journal articles.

## `filterType` (type: `string`):

Filter results to a specific scholarly work type. The default is journal articles.

## `sortOrder` (type: `string`):

The order in which results are returned. Published order sorts by publication date, relevance sorts by match to the search query.

## Actor input object example

```json
{
  "maxItems": 10,
  "filterType": "journal-article",
  "sortOrder": "published"
}
```

# Actor output Schema

## `results` (type: `string`):

Complete dataset of all scraped records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10,
    "filterType": "journal-article",
    "sortOrder": "published"
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/crossref-academic-paper-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "maxItems": 10,
    "filterType": "journal-article",
    "sortOrder": "published",
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/crossref-academic-paper-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10,
  "filterType": "journal-article",
  "sortOrder": "published"
}' |
apify call parseforge/crossref-academic-paper-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/crossref-academic-paper-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/X6zH2NC1U66zkMkpy/builds/zTS1dCX2OZbep6ZKh/openapi.json
