# OpenAlex Works Scraper (`neuton/openalex-works-scraper`) Actor

Search OpenAlex works by query, DOI, or OpenAlex ID. Extract scholarly titles, authorships, concepts, institutions, venues, publication dates, citations, and open-access metadata.

- **URL**: https://apify.com/neuton/openalex-works-scraper.md
- **Developed by:** [Ashwin Prasad](https://apify.com/neuton) (community)
- **Categories:** Education
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-usage

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## OpenAlex Works Scraper

Extract scholarly metadata from the OpenAlex API for academic intelligence, literature monitoring, citation analysis, and enrichment pipelines.

Use this actor to turn research queries, DOIs, or OpenAlex IDs into structured publication rows with authorship, institution, citation, concept, and open-access metadata. It is useful for universities, libraries, research products, AI/RAG builders, grants teams, and analysts mapping scientific topics or institutions.

### What it returns

- Work ID, DOI, title, publication year/date, type
- Authors, institutions, countries, venue/source
- Citation counts, referenced works, related works
- Open access status and landing/PDF URLs
- Concepts and keywords for research clustering

### Common use cases

- Build literature-review datasets with concepts, citations, authors, and institutions
- Monitor research output by topic, university, company, funder, or country
- Enrich DOI or title lists with OpenAlex IDs and open-access links
- Create RAG/semantic-search corpora from structured scholarly metadata
- Feed market maps for academic fields, biotech topics, AI labs, and grant scouting

### Input

Use search queries, DOIs, OpenAlex work IDs, or a mix of all three.

```json
{
  "queries": ["foundation models biology"],
  "dois": ["10.48550/arXiv.2303.08774"],
  "maxResultsPerQuery": 25,
  "includeReferencedWorks": true
}
```

### Output

Every result is saved to the default Apify dataset. Rows typically include `openalex_id`, `doi`, `title`, `publication_year`, `publication_date`, `type`, `authors`, `institutions`, `countries`, `source`, `cited_by_count`, `referenced_works`, `related_works`, `open_access`, `landing_page_url`, `pdf_url`, and concepts.

### SEO keywords

OpenAlex scraper, OpenAlex works API, scholarly works scraper, academic citation scraper, research metadata export, open access paper scraper, DOI enrichment API, scientific literature dataset.

### Pricing recommendation

Launch around $2 per 1,000 work rows. Keep launch pricing accessible for academic and AI-dataset users while preserving margin because the actor uses public APIs and does not require browser/proxy spend.

### Responsible use

This actor extracts public OpenAlex metadata. Respect source licenses and verify critical records before formal publication, grant decisions, procurement, or regulated analysis. Open-access links may point to third-party pages whose own terms still apply.

### Automation ideas

Schedule recurring runs for topics, institutions, authors, or funders. AI agents can summarize new research, cluster concepts, flag highly cited works, enrich grant pipelines, and update RAG indexes or market maps with fresh scholarly metadata.

# Actor input Schema

## `queries` (type: `array`):

Search queries

## `dois` (type: `array`):

DOIs

## `openAlexIds` (type: `array`):

OpenAlex IDs

## `maxResults` (type: `integer`):

Max results per query

## `email` (type: `string`):

Optional email passed to OpenAlex mailto parameter for better service routing.

## Actor input object example

```json
{
  "maxResults": 50
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("neuton/openalex-works-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("neuton/openalex-works-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call neuton/openalex-works-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=neuton/openalex-works-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/l4XfB7MEvC1eZ00MQ/builds/2XgCA0ZffPX4NhwTy/openapi.json
