OpenAlex Academic Research Scraper - Scholarly Papers avatar

OpenAlex Academic Research Scraper - Scholarly Papers

Pricing

from $2.00 / 1,000 results

Go to Apify Store
OpenAlex Academic Research Scraper - Scholarly Papers

OpenAlex Academic Research Scraper - Scholarly Papers

Search and extract academic papers, authors, institutions, and research topics from OpenAlex. Free open API covering 250M+ scholarly works. Get citations, abstracts, open access URLs.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

cloud9

cloud9

Maintained by Community

Actor stats

1

Bookmarked

3

Total users

0

Monthly active users

20 days ago

Last modified

Categories

Share

Search and extract academic papers, authors, institutions, and research topics from OpenAlex. Free open API covering 250M+ scholarly works. Get citations, abstracts, open access URLs.

Use cases

  • Map a research field's citation network and open-access share
  • Evaluate an author's or institution's output
  • Build a bibliometrics dashboard
  • Find open-access full-text links for a reading list
  • Feed a research RAG pipeline with abstracts

Input

ParameterTypeRequiredDefaultDescription
modestringYes"searchWorks"What to search for Allowed: searchWorks, searchAuthors, searchInstitutions, searchTopics.
searchQuerystringYes"machine learning"Search term
filterYearintegerNo—Filter by publication year (optional)
filterOpenAccessbooleanNofalseOnly return open access papers
sortBystringNo"relevance"Sort order for results Allowed: relevance, cited_by_count, publication_date.
maxResultsintegerNo50Maximum number of results
emailstringNo—Your email for polite pool (faster rate limits). Optional but recommended.

Example input

{
"mode": "searchWorks",
"searchQuery": "machine learning",
"filterOpenAccess": false,
"sortBy": "relevance",
"maxResults": 5
}

Output

The exact fields depend on the mode you run. This is real output from an actual run of this Actor:

{
"id": "https://openalex.org/W2101234009",
"doi": "https://doi.org/10.48550/arxiv.1201.0490",
"title": "Scikit-learn: Machine Learning in Python",
"authorNames": "Fabián Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Müller, Andreas, Nothman, Joe…",
"authors": [
{
"name": "Fabián Pedregosa",
"institution": "Commissariat à l'Énergie Atomique et aux Énergies Alternatives"
},
{
"name": "Gaël Varoquaux",
"institution": "Commissariat à l'Énergie Atomique et aux Énergies Alternatives"
}
],
"publicationYear": 2012,
"publicationDate": "2012-01-02",
"journal": "ORBi (University of Liège)",
"citedByCount": 64029,
"isOpenAccess": true,
"openAccessUrl": "https://orbi.uliege.be/handle/2268/225787",
"abstract": "Scikit-learn is a Python module integrating a wide range of state-of-the-art machine learning algorithms for medium-scale supervised and unsupervised …",
"concepts": [
"Python (programming language)",
"Documentation"
],
"type": "article"
}
FieldType
idstring
doistring
titlestring
authorNamesstring
authorsarray of object — each with name, institution
publicationYearnumber
publicationDatestring
journalstring
citedByCountnumber
isOpenAccessboolean
openAccessUrlstring
abstractstring
conceptsarray
typestring

The dataset also ships a preset table view (Overview), so the key columns are readable straight away in Apify Console, and exportable to JSON, CSV, Excel, or XML.

How to run it

In Apify Console — open the Actor, fill in the input form, click Start, then download the results from the Dataset tab.

With the JavaScript client

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('cloud9_ai/openalex-scraper').call({
"mode": "searchWorks",
"searchQuery": "machine learning",
"filterOpenAccess": false,
"sortBy": "relevance",
"maxResults": 5
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

With the Python client

from apify_client import ApifyClient
client = ApifyClient('YOUR_APIFY_TOKEN')
run = client.actor('cloud9_ai/openalex-scraper').call(run_input={
"mode": "searchWorks",
"searchQuery": "machine learning",
"filterOpenAccess": False,
"sortBy": "relevance",
"maxResults": 5
})
for item in client.dataset(run['defaultDatasetId']).iterate_items():
print(item)

With the API — POST https://api.apify.com/v2/acts/cloud9_ai~openalex-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN with the input JSON as the body.

Notes and limits

  • No API key, account, or login is needed — just the input above.
  • maxResults caps how much a single run collects, which is also what caps the run's cost.
  • Requests are paced and failed requests are retried automatically, so runs stay inside the source's rate limits.
  • Only publicly available data is collected. How you use the output is your responsibility, including the source's terms of use and any applicable data-protection law.

Support

Found a bug, or need a field that isn't in the output? Open an issue on the Issues tab of this Actor in Apify Console. Issues there are read and answered.

License

Apache-2.0