OpenAlex Scraper — Papers, Authors & Institutions avatar

OpenAlex Scraper — Papers, Authors & Institutions

Pricing

from $2.00 / 1,000 results

Go to Apify Store
OpenAlex Scraper — Papers, Authors & Institutions

OpenAlex Scraper — Papers, Authors & Institutions

Scrape OpenAlex scholarly data: research papers (title, authors, DOI, citations, venue, abstract, open access), authors (ORCID, affiliation, h-index, topics) and institutions. Search any entity with filters. For research, bibliometrics and researcher lead gen. Not affiliated with OpenAlex.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Haketa

Haketa

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

Search and scrape the open index of science: research papers (title, authors, DOI, citations, venue, open access, reconstructed abstract), authors (ORCID, affiliation, h-index, topics) and institutions — all from OpenAlex. Search any entity with powerful filters. Clean JSON/CSV/Excel in seconds. Built for research, bibliometrics, competitive science intelligence and researcher lead generation.


What This Actor Does

Pick an entity type and search the OpenAlex scholarly graph. You get one clean record per result:

Works (research papers)

Title, DOI, publication year/date, type, citation count, authors and their institutions, venue/journal, open-access status + URL, topics, language, and a reconstructed plain-text abstract.

Authors (researchers)

Name, ORCID, works count, citations, h-index and i10-index, last-known institution (+ country), affiliations and research topics — a ready researcher profile.

Institutions

Name, country, type, works count, citations, homepage, ROR ID, city/region and top topics.

Use search terms and/or OpenAlex filters (year, open access, country, and more).


Why Use This

  • The whole graph of science, free. 250M+ works, 90M+ authors and 100K+ institutions — no key, no login, no anti-bot.
  • Researcher lead generation. Author records come with ORCID, affiliation and h-index — ideal for outreach, recruiting and expert discovery.
  • Bibliometrics & intelligence. Citation counts, venues, topics and open-access status for any field or institution.
  • Clean abstracts. Abstracts are reconstructed into plain text (OpenAlex stores them inverted).

Quick Start

Run it in the console (no code)

  1. Choose an entity type: Works, Authors or Institutions.
  2. Add search terms and/or filters.
  3. Set Max results, click Start, export as JSON, CSV, Excel or HTML.

Pull papers in a field (Python)

from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run_input = {"entityType": "works", "searchTerms": ["machine learning"],
"filters": ["publication_year:2023", "is_oa:true"], "maxItems": 1000}
run = client.actor("YOUR_USERNAME/openalex-scraper").call(run_input=run_input)
for w in client.dataset(run["defaultDatasetId"]).iterate_items():
print(w["publicationYear"], w["citedByCount"], "·", w["title"], "·", w["doi"])

Build a researcher list (Python)

run = client.actor("YOUR_USERNAME/openalex-scraper").call(run_input={
"entityType": "authors", "searchTerms": ["Yann LeCun", "Yoshua Bengio"], "maxItems": 200,
})
for a in client.dataset(run["defaultDatasetId"]).iterate_items():
print(a["name"], "·", a["orcid"], "·", f"h={a['hIndex']}", "·", a["lastKnownInstitution"])

Input Parameters

FieldTypeDescription
entityTypestringworks, authors or institutions.
searchTermsarrayKeywords to search. See the note below on how search works per entity.
filtersarrayRaw OpenAlex filters, e.g. publication_year:2023, is_oa:true, authorships.institutions.country_code:us.
emailstringOptional email for OpenAlex's faster "polite pool" (recommended for large runs).
reconstructAbstractbooleanRebuild plain-text abstracts for works (default on).
maxItemsintegerMax results across all searches. 0 = no limit.
maxPagesPerSearchintegerPagination cap per search (200 per page).
proxyConfigurationobjectApify Proxy. Datacenter is enough (public API).

How search works per entity: for works and institutions, searchTerms matches titles/names and content (great for topic/keyword discovery). For authors, searchTerms matches the author name — to find researchers by topic instead, scrape works and read their authors, or use filters.


Output (works example)

{
"id": "W2100837269",
"title": "Scikit-learn: Machine Learning in Python",
"doi": "https://doi.org/10.48550/arxiv.1201.0490",
"publicationYear": 2012, "type": "article", "citedByCount": 63567,
"authorNames": ["Fabián Pedregosa", "Gaël Varoquaux", "…"],
"authorInstitutions": ["CEA"],
"venue": "Journal of Machine Learning Research",
"isOpenAccess": true, "oaUrl": "https://…",
"topics": ["Machine Learning", "Python Applications"],
"abstract": "Scikit-learn is a Python module integrating a wide range of …",
"openAlexUrl": "https://openalex.org/W2100837269"
}

About coverage: identity fields (title/name, id, citation and works counts) are present for essentially every record. DOI, abstract, venue and ORCID are present where OpenAlex has them — not every work has a DOI or abstract, and not every author has a registered ORCID. This reflects the source data, not a gap in scraping.


Use Cases

1. Research & literature review

Pull all papers on a topic with citations, venues, open-access links and abstracts.

2. Researcher discovery & lead generation

Build researcher lists with ORCID, affiliation and h-index for recruiting, outreach and expert networks.

3. Bibliometrics & competitive science intelligence

Analyse citation trends, institutions, topics and open-access rates across a field.

4. Institution benchmarking

Compare universities and labs by output, citations and research topics.


Tips

  • filters are powerful — combine with search, e.g. publication_year:2020-2024, is_oa:true, cited_by_count:>100.
  • For topic-based researcher discovery, scrape works and collect their authors (author search is name-based).
  • Add your email for OpenAlex's polite pool — faster and more consistent on large runs.
  • citedByCount and hIndex make ranking and filtering easy.
  • Schedule it with Apify Schedules to track new papers in your field.

Frequently Asked Questions

Do I need an account or key? No. OpenAlex is fully open — no login, key or anti-bot.

How do I find researchers by topic, not name? Author search matches names. To find researchers in a topic, scrape works for that topic and read their authors, or use filters.

Are abstracts included? Yes — for works, abstracts are reconstructed into plain text (OpenAlex stores them as an inverted index). Not every work has one.

What filters can I use? Any OpenAlex filter, e.g. publication_year, is_oa, type, authorships.institutions.country_code, cited_by_count.

What export formats are supported? JSON, CSV, Excel, HTML, or via API — plus Google Sheets, webhooks, Make and Zapier.


This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by OpenAlex or OurResearch. It reads only public, open scholarly data. Use the data responsibly and in line with applicable terms and laws.