OpenAlex Scraper
Pricing
Pay per usage
OpenAlex Scraper
Extract scholarly data from OpenAlex—titles, authors, institutions, venues, concepts—using this fast Apify actor. Get academic research in bulk via API, and export results as CSV, Excel, or HTML datasets for research, analytics, or discovery.
Pricing
Pay per usage
Rating
5.0
(1)
Developer
Shahid Irfan
Maintained by CommunityActor stats
0
Bookmarked
22
Total users
3
Monthly active users
3 days ago
Last modified
Categories
Share
What does OpenAlex Scholarly Data Scraper do?
OpenAlex Scholarly Data Scraper collects research-paper metadata, author profiles, institutional records, and publication-source details from OpenAlex. Choose an entity, enter a search term, and set a result limit to create a dataset for literature discovery, bibliometric analysis, or research monitoring.
For research works, the output includes titles, author names, affiliations, publication years, DOIs, citation counts, and abstracts when available. Download the results as JSON, CSV, Excel, or XML, or connect them to an Apify integration.
Why use OpenAlex Scholarly Data Scraper?
- Literature discovery - Collect paper titles, publication dates, DOIs, and available abstracts for an initial review of a research topic.
- Citation analysis - Sort results by citation count to identify frequently cited works or profiles within your search results.
- Researcher and institution lookup - Find named authors and organizations without manually copying profile details.
- Repeatable collection - Reuse the same inputs for scheduled runs and compare downloaded datasets over time.
- Batch output - Each completed page is saved during the run, so earlier results remain available if a later request fails.
- Bounded recovery - Temporary failures receive a limited number of retries rather than keeping the run waiting indefinitely.
The actor does not guarantee uninterrupted access, complete abstracts, or a particular number of matching records. Results depend on OpenAlex coverage and the available anonymous-access budget.
What data can you extract from OpenAlex?
| Entity | Data collected | Search guidance |
|---|---|---|
works | Paper titles, author names, affiliations, publication years, DOIs, abstracts, concepts, and citation counts | Use a research term or paper title |
authors | Display names, publication counts, citation counts, one last-known institution, and ORCID | Use an author's name; this is not a subject-expertise filter |
institutions | Organization names, countries, institution types, publication counts, and citation counts | Use an institution name |
sources | Journal or other source names, publisher/host organization names, ISSNs, counts, and open-access indicators | Use a publication-source name |
concepts | Concept names, levels, descriptions, publication counts, and citation counts | Legacy subject classification; OpenAlex has deprecated concepts |
The input also retains venues as a legacy option. Use sources for new publication-source searches; the actor sends the selected entity as supplied and does not translate legacy endpoint names. Topics, standalone publisher records, full-text downloads, and advanced geographic or date filters are not exposed by this actor.
How to use the actor
- Open the actor's input form in Apify Console.
- Select an entity, such as
worksfor research papers. - Enter a search term appropriate to that entity.
- Set
results_wantedand, for larger collections,max_pages. - Choose relevance or citation sorting, then start the run.
- Review the dataset preview and download your preferred format.
Start with 20 results to check whether your query produces useful records. The Console shows machine learning as a search example. Omitted or blank search input lists records without a keyword filter; it does not automatically run that example search.
Input Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
search | String | No | Empty | Research term, paper title, author name, institution name, or source name, depending on the selected entity |
entity | String | No | "works" | One of works, authors, institutions, sources, concepts, or the legacy venues option |
results_wanted | Integer | No | 20 | Maximum number of records to save; minimum 1 |
max_pages | Integer | No | 10 | Maximum number of result pages to request; minimum 1 |
sort | String | No | "relevance_score:desc" | Sort expression; use cited_by_count:desc for descending citation count. Relevance sorting is omitted when search is blank. |
A page contains at most 100 records. Both result and page limits apply: results_wanted: 500 with max_pages: 2 can collect at most 200 records, and fewer when matches or access are limited. OpenAlex's basic paging limit is 10,000 results; this actor does not provide cursor continuation beyond that limit.
There is no API-key or proxy input in the current version. Adding unrecognized credential or proxy properties to the input does not enable those features.
Output Data
Fields depend on the selected entity. Missing scalar metadata may be null; missing names or classifications may be empty arrays. Citation and publication counts describe the source data at collection time and can change later.
| Field | Type | Applies to | Description |
|---|---|---|---|
id | String or null | All | OpenAlex entity identifier |
url | String or null | All | Entity URL derived from its identifier |
source | String | All | Always openalex.org |
title | String or null | Works | Work title |
authors | Array | Works | Author display names |
institutions | Array | Works | Institution names from authorships; duplicates can occur |
publication_year | Integer or null | Works | Publication year |
doi | String or null | Works | DOI URL when available |
abstract | String or null | Works | Available abstract reconstructed as readable text; not the full paper |
concepts | Array | Works | Legacy concept display names |
cited_by_count | Integer | All | Citation count |
type | String or null | Works, institutions, sources/venues | Record type as supplied by OpenAlex |
display_name | String or null | Non-work entities | Entity name |
works_count | Integer | Non-work entities | Associated work count |
last_known_institution | String or null | Authors | First last-known institution when several are present |
orcid | String or null | Authors | Researcher's ORCID identifier |
country_code | String or null | Institutions | Country code |
publisher | String or null | Sources/venues | Host organization name, or publisher value when available |
issn_l | String or null | Sources/venues | Linking ISSN |
issn | Array or null | Sources/venues | ISSNs, with linking ISSN as a fallback when available |
is_oa | Boolean | Sources/venues | Open-access indicator |
is_in_doaj | Boolean | Sources/venues | Directory of Open Access Journals indicator |
level | Integer or null | Concepts | Legacy hierarchy level; the current mapping returns null for a zero level |
description | String or null | Concepts | Available concept description |
Usage Examples
Search research papers
Collect up to 20 works related to machine learning, ordered by relevance.
{"search": "machine learning","entity": "works","results_wanted": 20}
Look up an author
Search author names and rank matching profiles by citation count.
{"search": "Yoshua Bengio","entity": "authors","sort": "cited_by_count:desc","results_wanted": 5}
Find an institution
Collect matching institutional records for Stanford University.
{"search": "Stanford University","entity": "institutions","results_wanted": 5}
Collect a larger paper dataset
Request up to 300 climate-change works across at most three pages, ranked by citation count.
{"search": "climate change","entity": "works","sort": "cited_by_count:desc","results_wanted": 300,"max_pages": 3}
Sample Output
This work example uses metadata checked on September 30, 2026. Citation counts and affiliations may change. The abstract was unavailable for this record, so its value is null.
{"id": "https://openalex.org/W2919115771","url": "https://openalex.org/W2919115771","source": "openalex.org","title": "Deep learning","authors": ["Yann LeCun","Yoshua Bengio","Geoffrey E. Hinton"],"institutions": ["Meta (United States)","New York University","Université de Montréal","Google (United States)","University of Toronto"],"publication_year": 2015,"doi": "https://doi.org/10.1038/nature14539","abstract": null,"concepts": ["Computer science","Deep learning","Artificial intelligence","Abstraction","Representation (politics)","Layer (electronics)","Object (grammar)","Backpropagation","Convolutional neural network","Feature learning","Pattern recognition (psychology)","Speech recognition","Artificial neural network","Epistemology","Politics","Philosophy","Political science","Organic chemistry","Chemistry","Law"],"cited_by_count": 85092,"type": "article"}
Tips for Best Results
- Match the query to the entity - Use topics for works, people’s names for authors, and organization names for institutions.
- Check small runs first - Inspect 20 results before requesting hundreds of records.
- Choose the ranking deliberately - Citation sorting emphasizes frequently cited records, not necessarily the newest or most relevant research.
- Allow enough pages - Each page supplies at most 100 records; a low page limit can stop collection before the result target.
- Plan for missing data - An absent abstract, DOI, ORCID, or affiliation is not necessarily an extraction failure.
- Avoid overlapping large runs - Repeated anonymous requests can exhaust the available budget. Delay subsequent runs when OpenAlex reports a limit.
- Keep your earlier results - A failed later page does not remove batches already saved to the dataset.
Integrations
Use Apify's integrations and dataset access to connect the results to your workflow:
- Google Sheets or Airtable - Review citations, authors, and institutions in a spreadsheet or database.
- Make or Zapier - Connect completed runs to downstream research workflows.
- Slack - Send notifications through a configured integration.
- Webhooks - Trigger processing when a run completes or fails.
- Programmatic dataset access - Retrieve results for notebooks, dashboards, or a research database.
Export Formats
Download datasets as JSON, CSV, Excel, or XML. JSON preserves arrays such as author and institution lists for applications that need the original record structure.
Frequently Asked Questions
Is an API key required?
No. The current actor uses keyless access and does not accept an API key. Anonymous access has a limited daily budget, so a run can be throttled even when its input is valid.
What happens when OpenAlex returns a rate-limit error?
The actor makes at most four attempts per page: the initial request and three retries. It follows valid server-provided retry timing or uses short increasing waits with a small random variation when no timing is provided.
The run fails when retries are exhausted, a rate-limit response explicitly reports no remaining daily budget, or the requested cooldown exceeds 60 seconds. Malformed responses and permanent access or input errors are not retried. Wait for the limit to reset before starting another run; completed batches remain in the dataset.
Will a residential proxy solve HTTP 429 errors?
Not necessarily. A residential proxy may help when throttling is specific to the current network address, but it is not a guaranteed fix for an exhausted upstream budget or a service-wide restriction. It also adds cost and may increase response time.
This actor currently has no proxy integration. A proxy has not been tested against the reported cloud failure, so it is not presented as a verified solution. Respect OpenAlex's usage limits rather than relying on address rotation to avoid them.
Can I collect journals and publication venues?
Yes. Use entity: "sources" to search publication-source records. The publisher output describes the available host organization or publisher name; it is not a separate publisher profile. Prefer sources over the legacy venues input.
Are abstracts the full text of papers?
No. The actor returns an abstract only when OpenAlex supplies one. It does not download the article's full text or create a summary when an abstract is missing.
Can I find researchers by their expertise?
Not directly. Author searches match names. Searching authors for a subject such as artificial intelligence does not provide an expertise filter. Search works by subject and inspect their author names for an initial literature-based workflow.
How many records can I collect?
Collection stops at your result target, page limit, end of matching results, or an unrecoverable error. Each page contains at most 100 records, and basic paging cannot go beyond OpenAlex's 10,000-result limit. A requested count is a maximum, not a guarantee.
Can I schedule recurring runs?
Yes. Configure an Apify schedule with the desired inputs. Allow enough time between runs to avoid repeatedly consuming the anonymous budget, and monitor failures before scheduling larger collections.
What should I do if fields are empty?
Check several records and the entity type. Some fields apply only to works or profiles, and OpenAlex may not publish a value for a particular record. Missing metadata is represented by null values, empty arrays, or the mapped count/boolean defaults.
Support
Use the actor's Issues tab or contact the developer through Apify Console. Include the run URL, selected entity, result and page limits, and the relevant error message. Do not share private credentials.
Resources
Legal Notice
This actor collects publicly available scholarly metadata from OpenAlex. You are responsible for complying with OpenAlex's terms, applicable laws, and any requirements of your downstream workflow. Respect usage limits and handle author or institutional information responsibly. Metadata availability does not imply permission to download or redistribute the full text of a publication.