OpenAlex Academic Research Scraper - Scholarly Papers
Pricing
from $2.00 / 1,000 results
OpenAlex Academic Research Scraper - Scholarly Papers
Search and extract academic papers, authors, institutions, and research topics from OpenAlex. Free open API covering 250M+ scholarly works. Get citations, abstracts, open access URLs.
Pricing
from $2.00 / 1,000 results
Rating
0.0
(0)
Developer
cloud9
Maintained by CommunityActor stats
1
Bookmarked
3
Total users
0
Monthly active users
20 days ago
Last modified
Categories
Share
Search and extract academic papers, authors, institutions, and research topics from OpenAlex. Free open API covering 250M+ scholarly works. Get citations, abstracts, open access URLs.
Use cases
- Map a research field's citation network and open-access share
- Evaluate an author's or institution's output
- Build a bibliometrics dashboard
- Find open-access full-text links for a reading list
- Feed a research RAG pipeline with abstracts
Input
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
mode | string | Yes | "searchWorks" | What to search for Allowed: searchWorks, searchAuthors, searchInstitutions, searchTopics. |
searchQuery | string | Yes | "machine learning" | Search term |
filterYear | integer | No | — | Filter by publication year (optional) |
filterOpenAccess | boolean | No | false | Only return open access papers |
sortBy | string | No | "relevance" | Sort order for results Allowed: relevance, cited_by_count, publication_date. |
maxResults | integer | No | 50 | Maximum number of results |
email | string | No | — | Your email for polite pool (faster rate limits). Optional but recommended. |
Example input
{"mode": "searchWorks","searchQuery": "machine learning","filterOpenAccess": false,"sortBy": "relevance","maxResults": 5}
Output
The exact fields depend on the mode you run. This is real output from an actual run of this Actor:
{"id": "https://openalex.org/W2101234009","doi": "https://doi.org/10.48550/arxiv.1201.0490","title": "Scikit-learn: Machine Learning in Python","authorNames": "Fabián Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Müller, Andreas, Nothman, Joe…","authors": [{"name": "Fabián Pedregosa","institution": "Commissariat à l'Énergie Atomique et aux Énergies Alternatives"},{"name": "Gaël Varoquaux","institution": "Commissariat à l'Énergie Atomique et aux Énergies Alternatives"}],"publicationYear": 2012,"publicationDate": "2012-01-02","journal": "ORBi (University of Liège)","citedByCount": 64029,"isOpenAccess": true,"openAccessUrl": "https://orbi.uliege.be/handle/2268/225787","abstract": "Scikit-learn is a Python module integrating a wide range of state-of-the-art machine learning algorithms for medium-scale supervised and unsupervised …","concepts": ["Python (programming language)","Documentation"],"type": "article"}
| Field | Type |
|---|---|
id | string |
doi | string |
title | string |
authorNames | string |
authors | array of object — each with name, institution |
publicationYear | number |
publicationDate | string |
journal | string |
citedByCount | number |
isOpenAccess | boolean |
openAccessUrl | string |
abstract | string |
concepts | array |
type | string |
The dataset also ships a preset table view (Overview), so the key columns are readable straight away in Apify Console, and exportable to JSON, CSV, Excel, or XML.
How to run it
In Apify Console — open the Actor, fill in the input form, click Start, then download the results from the Dataset tab.
With the JavaScript client
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });const run = await client.actor('cloud9_ai/openalex-scraper').call({"mode": "searchWorks","searchQuery": "machine learning","filterOpenAccess": false,"sortBy": "relevance","maxResults": 5});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
With the Python client
from apify_client import ApifyClientclient = ApifyClient('YOUR_APIFY_TOKEN')run = client.actor('cloud9_ai/openalex-scraper').call(run_input={"mode": "searchWorks","searchQuery": "machine learning","filterOpenAccess": False,"sortBy": "relevance","maxResults": 5})for item in client.dataset(run['defaultDatasetId']).iterate_items():print(item)
With the API — POST https://api.apify.com/v2/acts/cloud9_ai~openalex-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN with the input JSON as the body.
Notes and limits
- No API key, account, or login is needed — just the input above.
maxResultscaps how much a single run collects, which is also what caps the run's cost.- Requests are paced and failed requests are retried automatically, so runs stay inside the source's rate limits.
- Only publicly available data is collected. How you use the output is your responsibility, including the source's terms of use and any applicable data-protection law.
Support
Found a bug, or need a field that isn't in the output? Open an issue on the Issues tab of this Actor in Apify Console. Issues there are read and answered.
License
Apache-2.0