OpenAlex Scraper | 20 Fields, Citations & OA, No API Key avatar

OpenAlex Scraper | 20 Fields, Citations & OA, No API Key

Pricing

from $0.60 / 1,000 work scrapeds

Go to Apify Store
OpenAlex Scraper | 20 Fields, Citations & OA, No API Key

OpenAlex Scraper | 20 Fields, Citations & OA, No API Key

Scrape 250M+ scholarly papers from OpenAlex as clean JSON. Filter by topic, year, citations, open access & type. Get authors, venues, abstracts. No API key. Use in Claude, ChatGPT & any MCP agent for literature reviews & RAG.

Pricing

from $0.60 / 1,000 work scrapeds

Rating

0.0

(0)

Developer

The Mine Works

The Mine Works

Maintained by Community

Actor stats

0

Bookmarked

16

Total users

5

Monthly active users

an hour ago

Last modified

Share

100 OpenAlex papers in 6 seconds

From The Mine Works, makers of Threads Scraper and B2B Leads Finder, with over 140,000 runs across 170+ public actors. This actor ranks #3 for "openalex scraper" in Apify Store search.

OpenAlex is the free, open index of the world's scholarly literature: hundreds of millions of papers, books, datasets and preprints with their authors, institutions, journals and citations. This actor searches it by topic and filters by year, citation count, work type and open access, then returns each work as one clean row with 19 fields (20 with the abstract): title, DOI, year, type, citation count, authors, their institutions, journal, open access status and free full text link, concepts, reference count and language. No API key and no pagination code.

Why choose this actor?

  • A literature shortlist in seconds. In our test of the live code on 2 October, 100 works on large language models published since 2023 came back in 6 seconds, out of 1,933,734 that matched. OpenAlex's results come 100 to a page, so even 1,000 works is only 10 requests.
  • The joined fields you would otherwise build yourself. Author names as a list, institutions deduplicated across all authors, the journal and its type, open access status with a direct link to the free PDF where one exists, concepts filtered to OpenAlex's confident ones, and the abstract rebuilt into readable text on request.
  • Cheap, and nothing for empty searches. $0.60 to $1.00 per 1,000 works depending on your plan, plus a $0.005 start fee per run. A search that matches nothing, failed requests and the summary row cost nothing beyond that.

Run it on Apify

Part of The Mine Works Science, health and government data family: CourtListener Scraper, Socrata Open Data Scraper, Academic Research MCP, FDA 510(k) Clearances Scraper, Crossref Scraper, PubMed Scraper.

Try it in one minute

Paste this into the JSON tab of the input page and press Start:

{
"searchTerm": "\"perovskite solar cells\"",
"fromYear": 2022,
"maxResults": 10
}

You get the 10 most cited works since 2022 that contain the exact phrase "perovskite solar cells", each with authors, institutions, journal and open access link, in a few seconds.

Everything is optional. Give a topic in searchTerm (words, or a phrase in double quotes for an exact match), narrow it with fromYear, toYear, minCitations, workType and openAccessOnly, and cap the run with maxResults. With no searchTerm at all, the filters alone pick a slice of the whole index, for example every review published in 2025 with at least 100 citations.

Apify's free plan includes $5 of credit every month, which covers about 4,900 works at this actor's Free plan price ($1.00 per 1,000 works, plus the $0.005 start fee on each run).

Copy to your AI assistant

themineworks/openalex-scholarly-works on Apify. Searches OpenAlex, the open index of scholarly works, and returns one row per work, most cited first, with title, DOI, year, date, type, citation count, authors, author institutions, journal and journal type, open access flag, status and free full text URL, concepts, reference count, language and optionally the abstract. Call ApifyClient("TOKEN").actor("themineworks/openalex-scholarly-works").call(run_input={...}), then client.dataset(run["defaultDatasetId"]).list_items().items. Required: nothing. Optional: searchTerm (string; put a phrase in double quotes for an exact match, because results are sorted by citations, not relevance), fromYear and toYear (integers), minCitations (integer), workType ("article", "review", "book-chapter", "book", "dataset", "preprint", "dissertation", "report" or "" for all; default ""), openAccessOnly (default false), includeAbstract (default false), maxResults (default 100, max 10000). The last row has _type "summary" with total_available and is never billed. Full spec: GET https://api.apify.com/v2/acts/themineworks~openalex-scholarly-works/builds/default (Bearer TOKEN), which returns inputSchema and readme. Token: https://console.apify.com/account/integrations?fpr=ymnoit&utm_source=apify-readme&utm_medium=referral

Key features

  • 19 fields per work, 20 with the abstract, including the authors list, deduplicated author_institutions, venue and venue_type, oa_status and oa_url, up to 8 concepts and referenced_works_count.
  • Five filters that combine freely: publication year range, minimum citation count, work type (8 types), open access only, and a full text search term.
  • Most cited first, always. Results are sorted by cited_by_count from highest to lowest, so a small maxResults gives you the most cited matching work. See the search tip below for getting relevant results out of this.
  • Readable abstracts. OpenAlex stores abstracts as a word position index; with includeAbstract on, the actor rebuilds them into normal text. In our test, 15 of 20 open access articles had one.
  • Up to 10,000 works per run, paged with OpenAlex's cursor 100 at a time, plus the total number of matching works in the summary row so you know how much is out there.
  • Official API, polite pool. The actor uses OpenAlex's own public API and identifies itself with a contact address, which puts its requests in OpenAlex's faster "polite pool". No key, no login, no scraping of web pages.

How to use it

Read this first: getting relevant results

OpenAlex's search looks for your words in titles, abstracts and full text, and this actor always sorts the matches by citation count. For a broad, unquoted term that means the most cited works in all of science that happen to contain the words come first. In a run on 30 September, large language models with no quotes returned R, lme4 and SHELX (statistics and crystallography software papers with tens of thousands of citations) as its top 3, and in a 100 work run only 15 titles mentioned language models.

Three habits fix this:

  1. Put the phrase in double quotes. "large language models" in quotes matched 6,707 works instead of 1,933,734, and the top results were about language models.
  2. Add a year range. fromYear keeps decades old classics out.
  3. Check the titles. Filter the rows on title or concepts before you use them, and look at total_available in the summary to see how wide your search was.

Basic: the most cited work on a topic

{
"searchTerm": "\"retrieval augmented generation\"",
"maxResults": 100
}

Recent, well cited and free to read

{
"searchTerm": "\"protein folding\"",
"fromYear": 2023,
"minCitations": 50,
"openAccessOnly": true,
"maxResults": 200
}

openAccessOnly keeps only works with a free full text, and oa_url points to it, so you can fetch and read every paper on the list.

Literature review shortlist with abstracts

{
"searchTerm": "\"gene editing\"",
"workType": "review",
"fromYear": 2020,
"toYear": 2026,
"includeAbstract": true,
"maxResults": 300
}

Review articles with their abstracts give you a map of a field in one run. Feed the abstracts to an embedding model or a classifier to cluster the topics.

Which labs and companies publish in a field

{
"searchTerm": "\"solid state battery\"",
"fromYear": 2022,
"maxResults": 1000
}

Count author_institutions across the rows to see who leads, and count concepts to see where the work is heading. Repeat monthly with a newer fromYear to watch it move.

A slice of the index with no search term

{
"fromYear": 2025,
"toYear": 2025,
"workType": "review",
"minCitations": 100,
"maxResults": 500
}

The filters alone select the works; the most cited come first.

Input parameters

ParameterTypeDefaultWhat it does
searchTermstringnone (form prefill: large language models)Words to find in titles, abstracts and full text. Use double quotes for an exact phrase. Leave empty to rely on the filters alone.
fromYearintegernoneOnly works published on or after 1 January of this year.
toYearintegernoneOnly works published on or before 31 December of this year.
minCitationsintegernoneOnly works cited at least this many times.
workTypestring"" (all types; form prefill: article)article, review, book-chapter, book, dataset, preprint, dissertation or report.
openAccessOnlybooleanfalseOnly works that are open access.
includeAbstractbooleanfalseRebuild and add each work's abstract. Rows get larger.
maxResultsinteger (1 to 10,000)100 (form prefill: 25)Most works to return, most cited first.

"Form prefill" values fill the Console form for you but are not defaults: an API call that leaves workType out searches all types. The run uses 512 MB of memory and a 300 second timeout by default; for runs near 10,000 works, a longer timeout is safer.

What data do you get?

One row per work. Fields OpenAlex has no value for are left out of that row rather than filled with null (for example venue when OpenAlex does not know where a work appeared, or oa_url for a closed work).

The work: openalex_id, openalex_url, doi, title, publication_year, publication_date, type (OpenAlex's own type, which can also be conference-paper and others beyond the eight filter values), language.

Who wrote it: authors (display names in author order), author_institutions (every institution named across all authors, deduplicated).

Where it appeared: venue (journal, conference or repository name), venue_type (such as journal or repository).

Impact and reach: cited_by_count, referenced_works_count (how many works it cites), is_open_access, oa_status (gold, green, hybrid, bronze, diamond or closed), oa_url (link to the free full text).

Topic: concepts (up to 8 concept labels that OpenAlex scores at 0.3 or above), abstract (with includeAbstract on, when OpenAlex has one).

Run data: scraped_at.

Each run ends with a _type: "summary" row: works_scraped, total_available (how many works in all of OpenAlex match your search and filters) and charged_for. A run that delivered works also adds a short _type: "info" note. Neither is a work and neither is ever charged.

About the data itself: the values are OpenAlex's, passed through as they are. OpenAlex builds its index automatically from many sources, and now and then a record carries an odd value, such as a citation count that belongs to a different work or a concept that does not fit. Check anything surprising against doi or openalex_url.

Stable fields for automations

These 14 fields were present in every work row we sampled (143 works from four runs, 30 September and 2 October 2026):

FieldWhat it holds
openalex_idOpenAlex work ID as a URL, such as https://openalex.org/W4323655724
titleTitle of the work
publication_yearYear, a number
publication_dateDate, YYYY-MM-DD
typeOpenAlex work type
cited_by_countCitations, a number
authorsArray of author names
author_institutionsArray of institution names
is_open_accesstrue or false
oa_statusOpenAlex open access status
conceptsArray of concept labels
referenced_works_countNumber of works it cites
openalex_urlLink to the work on OpenAlex
scraped_atISO timestamp when the row was captured

We will not rename these fields. New fields may be added over time; existing ones keep their names.

Output examples

Real rows from our test runs on 2 October 2026, with long lists and abstracts trimmed.

A journal article with a free full text link (large language models, from 2023):

{
"openalex_id": "https://openalex.org/W4323655724",
"doi": "https://doi.org/10.1016/j.lindif.2023.102274",
"title": "ChatGPT for good? On opportunities and challenges of large language models for education",
"publication_year": 2023,
"publication_date": "2023-03-09",
"type": "article",
"cited_by_count": 6795,
"authors": ["Enkelejda Kasneci", "Kathrin Seßler", "Stefan Küchemann", "Maria Bannert"],
"author_institutions": ["Technical University of Munich", "Ludwig-Maximilians-Universität München"],
"venue": "Learning and Individual Differences",
"venue_type": "journal",
"is_open_access": true,
"oa_status": "green",
"oa_url": "https://epub.ub.uni-muenchen.de/125071/1/ChatGPT_for_Good_v3.pdf",
"concepts": ["Curriculum", "Computer science", "Field (mathematics)", "Engineering ethics"],
"referenced_works_count": 75,
"language": "en",
"openalex_url": "https://openalex.org/W4323655724",
"scraped_at": "2026-10-02T15:25:53.435Z"
}

An open access article with its rebuilt abstract (perovskite solar, articles only, open access only, includeAbstract on):

{
"openalex_id": "https://openalex.org/W2301656337",
"doi": "https://doi.org/10.1039/c5ee03874j",
"title": "Cesium-containing triple cation perovskite solar cells: improved stability, reproducibility and high efficiency",
"publication_year": 2016,
"publication_date": "2016-01-01",
"type": "article",
"cited_by_count": 5470,
"authors": ["Michael Saliba", "Taisuke Matsui", "Ji-Youn Seo"],
"author_institutions": ["École Polytechnique Fédérale de Lausanne", "Panasonic (Japan)"],
"venue": "Energy & Environmental Science",
"venue_type": "journal",
"is_open_access": true,
"oa_status": "hybrid",
"oa_url": "https://pubs.rsc.org/en/content/articlepdf/2016/ee/c5ee03874j",
"concepts": ["Formamidinium", "Perovskite (structure)", "Caesium", "Photovoltaics"],
"referenced_works_count": 64,
"language": "en",
"openalex_url": "https://openalex.org/W2301656337",
"abstract": "Today's best perovskite solar cells use a mixture of formamidinium and methylammonium as the monovalent cations. With the addition of inorganic cesium, the resulting triple cation perovskite compositions are thermally more stable…",
"scraped_at": "2026-10-02T15:25:59.311Z"
}

A conference paper with no venue on record (the quoted search "large language models"):

{
"openalex_id": "https://openalex.org/W3133702157",
"doi": "https://doi.org/10.1145/3442188.3445922",
"title": "On the Dangers of Stochastic Parrots",
"publication_year": 2021,
"publication_date": "2021-03-01",
"type": "conference-paper",
"cited_by_count": 6730,
"authors": ["Emily M. Bender", "Timnit Gebru", "Angelina McMillan-Major", "Shmargaret Shmitchell"],
"author_institutions": ["University of Washington"],
"is_open_access": true,
"oa_status": "gold",
"oa_url": "https://doi.org/10.1145/3442188.3445922",
"concepts": ["Software deployment", "Computer science", "Stakeholder"],
"referenced_works_count": 117,
"language": "en",
"openalex_url": "https://openalex.org/W3133702157",
"scraped_at": "2026-10-02T15:27:03.953Z"
}

The summary row from the 100 work run:

{
"_type": "summary",
"works_scraped": 100,
"total_available": 1933734,
"charged_for": 0,
"scraped_at": "2026-10-02T15:25:53.487Z"
}

charged_for is 0 here only because this was a local test run, which Apify does not bill; on Apify it equals the number of works delivered.

Pricing

Pay per event: you pay for each work delivered to your dataset, plus a small start fee per run.

EventFreeBronzeSilverGold and above
work-scraped, per work$0.001$0.0009$0.00075$0.0006
work-scraped, per 1,000 works$1.00$0.90$0.75$0.60
apify-actor-start, per run$0.005 per GB of run memory, minimum one eventsamesamesame

The start fee, exactly. Apify's apify-actor-start event is charged once when a run starts, at $0.005 for each GB of memory the run uses, with a minimum of one event. This actor runs on 512 MB by default, so a default run pays the minimum, $0.005.

Worked examples on the Gold tier: 100 works cost $0.06 plus $0.005. 1,000 works cost $0.60 plus $0.005. A full 10,000 work run costs $6.00 plus $0.005. On the Free tier, 1,000 works cost $1.00 plus $0.005.

Never charged: a search that matches nothing, failed or rate limited requests, and the summary and info rows. A run that finds nothing pays only the start fee.

There is no scheduled price change for this actor. The Pricing tab on this page always shows the rate for your own plan; if it and this table ever differ, the Pricing tab is right.

FAQ

What is OpenAlex? OpenAlex is a free and open catalogue of scholarly works, authors, institutions, journals and citations, run by the nonprofit OurResearch as the open successor to Microsoft Academic Graph. Its data is released under CC0. This actor is independent and is not affiliated with OurResearch or OpenAlex.

Do I need an OpenAlex API key or an account? No. OpenAlex's API is free and needs no key. You need only an Apify account. The actor adds a contact address to its requests, which is how OpenAlex's faster "polite pool" works.

How are results ordered? By citation count, highest first. There is no relevance sort. That is why quoted phrases and a year range matter: see "Read this first: getting relevant results" above.

How many works can I get? Up to 10,000 per run. The summary row's total_available tells you how many works match in all; to go beyond 10,000, split the work by year range or work type.

How fresh is the data? Each run queries OpenAlex live. OpenAlex itself updates continuously, and citation counts change as new papers are indexed, so the same search can return slightly different numbers a week later.

Is the abstract always there? No. OpenAlex has abstracts for many works but not all; when one is missing, the abstract field is left out of that row. In our test, 15 of 20 open access articles had one.

What does oa_status mean? OpenAlex's open access category: gold (open in an open access journal), green (a free copy in a repository), hybrid (open article in a subscription journal), bronze (free to read on the publisher's site without an open licence), diamond (open access journal with no author fees) or closed. oa_url links to the free copy when there is one.

Why are there at most 8 concepts? OpenAlex tags each work with many concepts and a confidence score. The actor keeps those scoring 0.3 or above and caps the list at 8, because the long tail below that is mostly noise.

Can I run it on a schedule? Yes. Save your input as a task, then in Apify Console open Schedules, choose Create new and pick a time or a cron expression such as 0 6 1 * * (monthly). The actor does not remember earlier runs, so to track new papers set fromYear to the current year and remove repeats downstream by openalex_id.

How do I export the data? From the run's Storage tab as JSON, CSV, Excel, XML or HTML, or through the Apify API. Remove rows that have a _type field to keep works only. In CSV and Excel, Apify spreads array fields such as authors over numbered columns.

Can I use it from Claude, ChatGPT or another AI assistant?

  • Connector URL: https://mcp.apify.com/?tools=themineworks/openalex-scholarly-works.
  • Claude: Settings > Connectors > Add custom connector, paste the URL, sign in with Apify.
  • ChatGPT: developer mode, add an MCP connector with the URL, sign in with Apify.
  • Cursor or VS Code: add it as an HTTP MCP server with that URL.
  • Claude Code: claude mcp add -t http openalex-scholarly-works "https://mcp.apify.com/?tools=themineworks/openalex-scholarly-works".

Is it legal to use OpenAlex data? Yes, in general: OpenAlex publishes its metadata under the CC0 public domain dedication and offers the API for this kind of use. Author names are part of the public scholarly record. The full texts that oa_url points to carry their own licences, so check those before you republish a paper. You remain responsible for how you use the data, including data protection laws such as GDPR and CCPA. This is general information, not legal advice.

Integrations

  • Google Sheets: export a run to a sheet for a shareable reading list.
  • Make, Zapier and n8n: use the Apify app or node to run a monthly search and post new highly cited papers to Slack or Notion.
  • Webhooks: have Apify call your URL when a run succeeds, then fetch the dataset.
  • API and client libraries: start runs and read datasets from Python, JavaScript or any HTTP client. The "Copy to your AI assistant" block above has the exact call.
  • MCP clients: Claude, ChatGPT, Cursor, VS Code and other MCP clients can call the actor through https://mcp.apify.com.

More from The Mine Works

Science, health and government data

Social media and video

Leads and business directories

Marketing, SEO and reviews

LinkedIn

Real estate

Jobs and hiring

Company and business data

E-commerce and marketplaces

Food and local services

Developer and AI tools

More tools

Support

Found a bug or need a field we do not return yet? Open an issue on the Issues tab of this page and we will reply there. To ask for a new scholarly source or data set, email dmineworks@gmail.com.

OpenAlex Scraper turns OpenAlex searches into clean rows of works with authors, institutions, journals, citations and open access links, most cited first, from $0.60 per 1,000 works.