OpenAlex Scraper avatar

OpenAlex Scraper

Pricing

Pay per usage

Go to Apify Store
OpenAlex Scraper

OpenAlex Scraper

Extract scholarly data from OpenAlex—titles, authors, institutions, venues, concepts—using this fast Apify actor. Get academic research in bulk via API, and export results as CSV, Excel, or HTML datasets for research, analytics, or discovery.

Pricing

Pay per usage

Rating

5.0

(1)

Developer

Shahid Irfan

Shahid Irfan

Maintained by Community

Actor stats

0

Bookmarked

22

Total users

3

Monthly active users

3 days ago

Last modified

Share

What does OpenAlex Scholarly Data Scraper do?

OpenAlex Scholarly Data Scraper collects research-paper metadata, author profiles, institutional records, and publication-source details from OpenAlex. Choose an entity, enter a search term, and set a result limit to create a dataset for literature discovery, bibliometric analysis, or research monitoring.

For research works, the output includes titles, author names, affiliations, publication years, DOIs, citation counts, and abstracts when available. Download the results as JSON, CSV, Excel, or XML, or connect them to an Apify integration.

Why use OpenAlex Scholarly Data Scraper?

  • Literature discovery - Collect paper titles, publication dates, DOIs, and available abstracts for an initial review of a research topic.
  • Citation analysis - Sort results by citation count to identify frequently cited works or profiles within your search results.
  • Researcher and institution lookup - Find named authors and organizations without manually copying profile details.
  • Repeatable collection - Reuse the same inputs for scheduled runs and compare downloaded datasets over time.
  • Batch output - Each completed page is saved during the run, so earlier results remain available if a later request fails.
  • Bounded recovery - Temporary failures receive a limited number of retries rather than keeping the run waiting indefinitely.

The actor does not guarantee uninterrupted access, complete abstracts, or a particular number of matching records. Results depend on OpenAlex coverage and the available anonymous-access budget.

What data can you extract from OpenAlex?

EntityData collectedSearch guidance
worksPaper titles, author names, affiliations, publication years, DOIs, abstracts, concepts, and citation countsUse a research term or paper title
authorsDisplay names, publication counts, citation counts, one last-known institution, and ORCIDUse an author's name; this is not a subject-expertise filter
institutionsOrganization names, countries, institution types, publication counts, and citation countsUse an institution name
sourcesJournal or other source names, publisher/host organization names, ISSNs, counts, and open-access indicatorsUse a publication-source name
conceptsConcept names, levels, descriptions, publication counts, and citation countsLegacy subject classification; OpenAlex has deprecated concepts

The input also retains venues as a legacy option. Use sources for new publication-source searches; the actor sends the selected entity as supplied and does not translate legacy endpoint names. Topics, standalone publisher records, full-text downloads, and advanced geographic or date filters are not exposed by this actor.

How to use the actor

  1. Open the actor's input form in Apify Console.
  2. Select an entity, such as works for research papers.
  3. Enter a search term appropriate to that entity.
  4. Set results_wanted and, for larger collections, max_pages.
  5. Choose relevance or citation sorting, then start the run.
  6. Review the dataset preview and download your preferred format.

Start with 20 results to check whether your query produces useful records. The Console shows machine learning as a search example. Omitted or blank search input lists records without a keyword filter; it does not automatically run that example search.

Input Parameters

ParameterTypeRequiredDefaultDescription
searchStringNoEmptyResearch term, paper title, author name, institution name, or source name, depending on the selected entity
entityStringNo"works"One of works, authors, institutions, sources, concepts, or the legacy venues option
results_wantedIntegerNo20Maximum number of records to save; minimum 1
max_pagesIntegerNo10Maximum number of result pages to request; minimum 1
sortStringNo"relevance_score:desc"Sort expression; use cited_by_count:desc for descending citation count. Relevance sorting is omitted when search is blank.

A page contains at most 100 records. Both result and page limits apply: results_wanted: 500 with max_pages: 2 can collect at most 200 records, and fewer when matches or access are limited. OpenAlex's basic paging limit is 10,000 results; this actor does not provide cursor continuation beyond that limit.

There is no API-key or proxy input in the current version. Adding unrecognized credential or proxy properties to the input does not enable those features.

Output Data

Fields depend on the selected entity. Missing scalar metadata may be null; missing names or classifications may be empty arrays. Citation and publication counts describe the source data at collection time and can change later.

FieldTypeApplies toDescription
idString or nullAllOpenAlex entity identifier
urlString or nullAllEntity URL derived from its identifier
sourceStringAllAlways openalex.org
titleString or nullWorksWork title
authorsArrayWorksAuthor display names
institutionsArrayWorksInstitution names from authorships; duplicates can occur
publication_yearInteger or nullWorksPublication year
doiString or nullWorksDOI URL when available
abstractString or nullWorksAvailable abstract reconstructed as readable text; not the full paper
conceptsArrayWorksLegacy concept display names
cited_by_countIntegerAllCitation count
typeString or nullWorks, institutions, sources/venuesRecord type as supplied by OpenAlex
display_nameString or nullNon-work entitiesEntity name
works_countIntegerNon-work entitiesAssociated work count
last_known_institutionString or nullAuthorsFirst last-known institution when several are present
orcidString or nullAuthorsResearcher's ORCID identifier
country_codeString or nullInstitutionsCountry code
publisherString or nullSources/venuesHost organization name, or publisher value when available
issn_lString or nullSources/venuesLinking ISSN
issnArray or nullSources/venuesISSNs, with linking ISSN as a fallback when available
is_oaBooleanSources/venuesOpen-access indicator
is_in_doajBooleanSources/venuesDirectory of Open Access Journals indicator
levelInteger or nullConceptsLegacy hierarchy level; the current mapping returns null for a zero level
descriptionString or nullConceptsAvailable concept description

Usage Examples

Search research papers

Collect up to 20 works related to machine learning, ordered by relevance.

{
"search": "machine learning",
"entity": "works",
"results_wanted": 20
}

Look up an author

Search author names and rank matching profiles by citation count.

{
"search": "Yoshua Bengio",
"entity": "authors",
"sort": "cited_by_count:desc",
"results_wanted": 5
}

Find an institution

Collect matching institutional records for Stanford University.

{
"search": "Stanford University",
"entity": "institutions",
"results_wanted": 5
}

Collect a larger paper dataset

Request up to 300 climate-change works across at most three pages, ranked by citation count.

{
"search": "climate change",
"entity": "works",
"sort": "cited_by_count:desc",
"results_wanted": 300,
"max_pages": 3
}

Sample Output

This work example uses metadata checked on September 30, 2026. Citation counts and affiliations may change. The abstract was unavailable for this record, so its value is null.

{
"id": "https://openalex.org/W2919115771",
"url": "https://openalex.org/W2919115771",
"source": "openalex.org",
"title": "Deep learning",
"authors": [
"Yann LeCun",
"Yoshua Bengio",
"Geoffrey E. Hinton"
],
"institutions": [
"Meta (United States)",
"New York University",
"Université de Montréal",
"Google (United States)",
"University of Toronto"
],
"publication_year": 2015,
"doi": "https://doi.org/10.1038/nature14539",
"abstract": null,
"concepts": [
"Computer science",
"Deep learning",
"Artificial intelligence",
"Abstraction",
"Representation (politics)",
"Layer (electronics)",
"Object (grammar)",
"Backpropagation",
"Convolutional neural network",
"Feature learning",
"Pattern recognition (psychology)",
"Speech recognition",
"Artificial neural network",
"Epistemology",
"Politics",
"Philosophy",
"Political science",
"Organic chemistry",
"Chemistry",
"Law"
],
"cited_by_count": 85092,
"type": "article"
}

Tips for Best Results

  • Match the query to the entity - Use topics for works, people’s names for authors, and organization names for institutions.
  • Check small runs first - Inspect 20 results before requesting hundreds of records.
  • Choose the ranking deliberately - Citation sorting emphasizes frequently cited records, not necessarily the newest or most relevant research.
  • Allow enough pages - Each page supplies at most 100 records; a low page limit can stop collection before the result target.
  • Plan for missing data - An absent abstract, DOI, ORCID, or affiliation is not necessarily an extraction failure.
  • Avoid overlapping large runs - Repeated anonymous requests can exhaust the available budget. Delay subsequent runs when OpenAlex reports a limit.
  • Keep your earlier results - A failed later page does not remove batches already saved to the dataset.

Integrations

Use Apify's integrations and dataset access to connect the results to your workflow:

  • Google Sheets or Airtable - Review citations, authors, and institutions in a spreadsheet or database.
  • Make or Zapier - Connect completed runs to downstream research workflows.
  • Slack - Send notifications through a configured integration.
  • Webhooks - Trigger processing when a run completes or fails.
  • Programmatic dataset access - Retrieve results for notebooks, dashboards, or a research database.

Export Formats

Download datasets as JSON, CSV, Excel, or XML. JSON preserves arrays such as author and institution lists for applications that need the original record structure.

Frequently Asked Questions

Is an API key required?

No. The current actor uses keyless access and does not accept an API key. Anonymous access has a limited daily budget, so a run can be throttled even when its input is valid.

What happens when OpenAlex returns a rate-limit error?

The actor makes at most four attempts per page: the initial request and three retries. It follows valid server-provided retry timing or uses short increasing waits with a small random variation when no timing is provided.

The run fails when retries are exhausted, a rate-limit response explicitly reports no remaining daily budget, or the requested cooldown exceeds 60 seconds. Malformed responses and permanent access or input errors are not retried. Wait for the limit to reset before starting another run; completed batches remain in the dataset.

Will a residential proxy solve HTTP 429 errors?

Not necessarily. A residential proxy may help when throttling is specific to the current network address, but it is not a guaranteed fix for an exhausted upstream budget or a service-wide restriction. It also adds cost and may increase response time.

This actor currently has no proxy integration. A proxy has not been tested against the reported cloud failure, so it is not presented as a verified solution. Respect OpenAlex's usage limits rather than relying on address rotation to avoid them.

Can I collect journals and publication venues?

Yes. Use entity: "sources" to search publication-source records. The publisher output describes the available host organization or publisher name; it is not a separate publisher profile. Prefer sources over the legacy venues input.

Are abstracts the full text of papers?

No. The actor returns an abstract only when OpenAlex supplies one. It does not download the article's full text or create a summary when an abstract is missing.

Can I find researchers by their expertise?

Not directly. Author searches match names. Searching authors for a subject such as artificial intelligence does not provide an expertise filter. Search works by subject and inspect their author names for an initial literature-based workflow.

How many records can I collect?

Collection stops at your result target, page limit, end of matching results, or an unrecoverable error. Each page contains at most 100 records, and basic paging cannot go beyond OpenAlex's 10,000-result limit. A requested count is a maximum, not a guarantee.

Can I schedule recurring runs?

Yes. Configure an Apify schedule with the desired inputs. Allow enough time between runs to avoid repeatedly consuming the anonymous budget, and monitor failures before scheduling larger collections.

What should I do if fields are empty?

Check several records and the entity type. Some fields apply only to works or profiles, and OpenAlex may not publish a value for a particular record. Missing metadata is represented by null values, empty arrays, or the mapped count/boolean defaults.

Support

Use the actor's Issues tab or contact the developer through Apify Console. Include the run URL, selected entity, result and page limits, and the relevant error message. Do not share private credentials.

Resources

This actor collects publicly available scholarly metadata from OpenAlex. You are responsible for complying with OpenAlex's terms, applicable laws, and any requirements of your downstream workflow. Respect usage limits and handle author or institutional information responsibly. Metadata availability does not imply permission to download or redistribute the full text of a publication.