OpenAlex Scraper - 250M+ Papers, Abstracts & Citations avatar

OpenAlex Scraper - 250M+ Papers, Abstracts & Citations

Pricing

from $1.94 / 1,000 works

Go to Apify Store
OpenAlex Scraper - 250M+ Papers, Abstracts & Citations

OpenAlex Scraper - 250M+ Papers, Abstracts & Citations

Search OpenAlex papers by keyword and get the abstract back as readable text. Each row holds the title, authors, institutions, year, venue, citation count, topic tags, DOI and the open-access link when a free copy is known. Raw OpenAlex filters work too. $2.00 per 1,000 works.

Pricing

from $1.94 / 1,000 works

Rating

5.0

(1)

Developer

Dami's Studio

Dami's Studio

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

0

Monthly active users

2 days ago

Last modified

Share

Search OpenAlex by keyword and get one row per paper with the abstract as readable text: the title, the authors and their institutions, the year, the citation count, the DOI, whether it is open access and where to read it for free.

The abstract is the part that takes work. OpenAlex does not store abstracts as text. It stores a word-position index, and this actor rebuilds the sentences from it. A paper with no index has no abstract, and the field is null rather than a guess.

InputSearch keywords, with an optional OpenAlex filter
OutputOne row per work: title, abstract, authors, institutions, year and date, type, venue, citations, topics, DOI, open-access link
Ceiling10,000 works per run
Account neededNone, and no API key
Price$2.00 per 1,000 works, flat on every plan. The free plan's $5 a month covers about 2,500

🔍 What OpenAlex Scraper does

Your keywords go against OpenAlex's own search, which reads titles, abstracts and full text where it has them. The run then pages through the matches, fifty at a time, until it has the number of unique works you asked for or OpenAlex runs out.

You can order by relevance, by citation count or by date, and set a published-on-or-after floor. For anything more specific there is a raw filter field that goes straight through to OpenAlex, so type:article,is_oa:true works exactly as their documentation describes it.

Nothing is written to your dataset until the whole set has been collected, so you get either the full result set or a diagnostic row explaining why not.

📋 What data you get from each OpenAlex paper

What you getField
OpenAlex's permanent id, and the DOI when there is oneopenalexId, doi
Title and the abstract as plain texttitle, abstract
Authors, and the institutions they were affiliated withauthors, institutions
Year and full publication dateyear, publicationDate
What kind of work it is, and where it appearedtype, venue
How many works cite itcitations
OpenAlex's topic tagsconcepts
Whether a free copy is known, and whereisOpenAccess, oaUrl
A link to the work on OpenAlexurl

▶️ How to scrape OpenAlex papers

  1. Open OpenAlex Scraper and click Try for free.
  2. Type your keywords into Search query.
  3. Set Max works. Start around 50 to see the row shape.
  4. Pick a Sort by, add a From publication date if you want, then click Start.
  5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.

💰 How much does it cost to scrape OpenAlex?

$2.00 per 1,000 works. Flat on every Apify plan, no volume tiers. On the free plan, the $5 Apify gives you each month covers about 2,500 works.

You pay per work delivered. Works that appear twice across pages are dropped before they are counted, diagnostic rows are not charged, and a search that matches nothing is not charged. Set a maximum cost on the run and it stops when it gets there.

📥 What you give it

{
"query": "protein folding",
"sort": "citations",
"fromDate": "2020-01-01",
"maxItems": 500
}
FieldDefaultWhat it is
querynone, the form starts with machine learningThe keywords to search for. Required.
fromDatenoneYYYY-MM-DD. Only works published on or after this date.
sortrelevancerelevance, citations for most cited, or date for newest first.
filternoneAn OpenAlex filter string, passed through untouched. Comma-separated key:value pairs such as type:article,is_oa:true.
maxItems100How many works to return, up to 10,000.
notionConnectornoneOptional. Writes each work into your Notion once the run finishes.
notionParentIdnoneOptional. The Notion data source to write into.
proxyConfigurationoffOptional network settings. Off by default, and a normal run does not need it.

filter is handed to OpenAlex as you typed it, and a mistake in it does not come back as an error. OpenAlex answers a malformed filter with an empty result set, so the run finishes on NO_RESULTS and you are left thinking your query matched nothing. Check the filter keys against OpenAlex's documentation before blaming the query.

If your filter already contains from_publication_date, the fromDate field is ignored. Set the date in one place, not both.

📤 What you get back

A real row from a run on 2 October 2026, with the lists and the abstract cut short (the row carries 19 authors and 10 institutions):

{
"ok": true,
"openalexId": "https://openalex.org/W2101234009",
"doi": "https://doi.org/10.48550/arxiv.1201.0490",
"title": "Scikit-learn: Machine Learning in Python",
"authors": ["Fabián Pedregosa", "Gaël Varoquaux", "Alexandre Gramfort", "Michel, Vincent", "..."],
"institutions": ["Commissariat à l'Énergie Atomique et aux Énergies Alternatives", "..."],
"year": 2012,
"publicationDate": "2012-01-02",
"type": "article",
"venue": "ORBi (University of Liège)",
"citations": 63573,
"concepts": ["Python (programming language)", "Documentation", "Computer science", "MIT License", "Artificial intelligence"],
"isOpenAccess": true,
"oaUrl": "https://orbi.uliege.be/handle/2268/225787",
"abstract": "Scikit-learn is a Python module integrating a wide range of ... machine learning algorithms for medium-scale supervised and unsupervised problems. ...",
"url": "https://openalex.org/W2101234009"
}
FieldHow to read it
openalexIdOpenAlex's permanent id. Use it as your key, since not every work has a DOI.
abstractRebuilt from OpenAlex's word-position index. null where there is no index to rebuild from.
authorsNames as OpenAlex holds them, so a few arrive family name first, like Michel, Vincent.
venueOpenAlex's primary location for the work. Often the journal, but it can be a repository, as above, or null.
citationsOpenAlex's cited-by count, and 0 when OpenAlex has no figure at all.
institutionsThe affiliations OpenAlex resolved for the authors, without repeats. Sometimes messy, because the matching is theirs.
conceptsUp to five of OpenAlex's own topic tags, not keywords the authors chose.
isOpenAccess, oaUrlWhether a free copy is known, and where it is. oaUrl is null when none is known.
typearticle, preprint, book-chapter, dataset and so on, in OpenAlex's vocabulary.

🧾 Reading the output

Two kinds of row land in your dataset, and ok tells them apart.

RowHow to spot itBilled
A workok: true and an openalexIdyes
A diagnosticok: false and an errorCodeno

The overview table in the Apify Console shows the work columns only, so a diagnostic row looks blank there. Switch to the JSON or All fields view to read it.

CodeWhat it means
BAD_INPUTNo query was given.
NO_RESULTSThe request worked and nothing came back. A malformed filter also lands here.
RATE_LIMITEDOpenAlex asked for a slower pace than the run could keep. Try a smaller run.
SERVER_ERROROpenAlex answered with a server error. Usually passes.
BLOCKEDOpenAlex refused the request. Re-run it.
NETWORKOpenAlex was unreachable. Re-run it.

💡 What people use it for

  • Pulling a field's most-cited work with the abstracts attached, in one file, for a reading list.
  • Feeding titles and abstracts to a model that has to summarise research it was not trained on.
  • Finding which institutions keep appearing on a topic, using the institutions field.
  • Filtering to open access with is_oa:true, so every row has somewhere to read it for free.
  • Watching a topic on a schedule with sort on date, so each run surfaces what is new.

🚧 What it does not do

  • Keyword search over works only. No author, institution or funder lookup by entity.
  • No full text. You get the abstract and a link, not the paper.
  • Abstracts are missing where OpenAlex has no index for them, and that is common on older and paywalled records.
  • citations cannot tell you zero from unknown. Both read 0.
  • A bad filter looks like an empty search, not like an error.
  • No end date field. Use the raw filter if you need an upper bound.
  • Institutions and concepts are OpenAlex's matching, and it is not perfect. Treat them as hints, not as ground truth.
  • All or nothing. A failure partway through discards what had been collected, so you get a diagnostic row rather than a partial set.

🧭 Which research scraper do you need?

If you wantUse
arXiv preprints with full abstracts and PDF linksarXiv Scraper
Anything with a registered DOI, with its journal, publisher and ISSNCrossref Scraper
Papers with readable abstracts, institutions and open-access linksThis one
Patents rather than papersGoogle Patents Search Scraper
Book records by keyword, title or ISBNBooks Scraper
Encyclopedia articles as plain textWikipedia Scraper
arXiv, OpenAlex and Wikipedia searches as tools for an AI agentResearch MCP Server

❓ Questions people ask

Do I need an OpenAlex account or an API key?

No. OpenAlex is open and the actor reads it directly.

Why is the abstract null on some papers?

OpenAlex only keeps a word-position index for a work if the publisher made one available. No index, no abstract to rebuild.

Why is the venue a repository instead of a journal?

OpenAlex picks one primary location per work, and for a paper with a preprint or a repository copy it sometimes picks that. The DOI still points at the published version.

Can I search by author or institution?

Not as a mode. You can approximate it through the raw filter field, using OpenAlex's own filter keys.

Can I call it from code or connect it to an AI assistant?

Yes. The API tab has ready-made code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/openalex-scraper. Either way the run happens on your Apify account at the same price.

OpenAlex publishes this data openly under a public domain dedication and invites this kind of use. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the run ID, the query and any filter you used. The errorCode on the diagnostic row usually names the problem on its own.