OpenAlex Scraper - 250M+ Papers, Abstracts & Citations
Pricing
from $1.94 / 1,000 works
OpenAlex Scraper - 250M+ Papers, Abstracts & Citations
Search OpenAlex papers by keyword and get the abstract back as readable text. Each row holds the title, authors, institutions, year, venue, citation count, topic tags, DOI and the open-access link when a free copy is known. Raw OpenAlex filters work too. $2.00 per 1,000 works.
Pricing
from $1.94 / 1,000 works
Rating
5.0
(1)
Developer
Dami's Studio
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
0
Monthly active users
2 days ago
Last modified
Categories
Share
Search OpenAlex by keyword and get one row per paper with the abstract as readable text: the title, the authors and their institutions, the year, the citation count, the DOI, whether it is open access and where to read it for free.
The abstract is the part that takes work. OpenAlex does not store abstracts as text. It stores a
word-position index, and this actor rebuilds the sentences from it. A paper with no index has no
abstract, and the field is null rather than a guess.
| Input | Search keywords, with an optional OpenAlex filter |
| Output | One row per work: title, abstract, authors, institutions, year and date, type, venue, citations, topics, DOI, open-access link |
| Ceiling | 10,000 works per run |
| Account needed | None, and no API key |
| Price | $2.00 per 1,000 works, flat on every plan. The free plan's $5 a month covers about 2,500 |
🔍 What OpenAlex Scraper does
Your keywords go against OpenAlex's own search, which reads titles, abstracts and full text where it has them. The run then pages through the matches, fifty at a time, until it has the number of unique works you asked for or OpenAlex runs out.
You can order by relevance, by citation count or by date, and set a published-on-or-after floor. For
anything more specific there is a raw filter field that goes straight through to OpenAlex, so
type:article,is_oa:true works exactly as their documentation describes it.
Nothing is written to your dataset until the whole set has been collected, so you get either the full result set or a diagnostic row explaining why not.
📋 What data you get from each OpenAlex paper
| What you get | Field |
|---|---|
| OpenAlex's permanent id, and the DOI when there is one | openalexId, doi |
| Title and the abstract as plain text | title, abstract |
| Authors, and the institutions they were affiliated with | authors, institutions |
| Year and full publication date | year, publicationDate |
| What kind of work it is, and where it appeared | type, venue |
| How many works cite it | citations |
| OpenAlex's topic tags | concepts |
| Whether a free copy is known, and where | isOpenAccess, oaUrl |
| A link to the work on OpenAlex | url |
▶️ How to scrape OpenAlex papers
- Open OpenAlex Scraper and click Try for free.
- Type your keywords into Search query.
- Set Max works. Start around 50 to see the row shape.
- Pick a Sort by, add a From publication date if you want, then click Start.
- Download the dataset as JSON, CSV or Excel, or read it from the Apify API.
💰 How much does it cost to scrape OpenAlex?
$2.00 per 1,000 works. Flat on every Apify plan, no volume tiers. On the free plan, the $5 Apify gives you each month covers about 2,500 works.
You pay per work delivered. Works that appear twice across pages are dropped before they are counted, diagnostic rows are not charged, and a search that matches nothing is not charged. Set a maximum cost on the run and it stops when it gets there.
📥 What you give it
{"query": "protein folding","sort": "citations","fromDate": "2020-01-01","maxItems": 500}
| Field | Default | What it is |
|---|---|---|
query | none, the form starts with machine learning | The keywords to search for. Required. |
fromDate | none | YYYY-MM-DD. Only works published on or after this date. |
sort | relevance | relevance, citations for most cited, or date for newest first. |
filter | none | An OpenAlex filter string, passed through untouched. Comma-separated key:value pairs such as type:article,is_oa:true. |
maxItems | 100 | How many works to return, up to 10,000. |
notionConnector | none | Optional. Writes each work into your Notion once the run finishes. |
notionParentId | none | Optional. The Notion data source to write into. |
proxyConfiguration | off | Optional network settings. Off by default, and a normal run does not need it. |
filter is handed to OpenAlex as you typed it, and a mistake in it does not come back as an
error. OpenAlex answers a malformed filter with an empty result set, so the run finishes on
NO_RESULTS and you are left thinking your query matched nothing. Check the filter keys against
OpenAlex's documentation before blaming the query.
If your filter already contains from_publication_date, the fromDate field is ignored. Set
the date in one place, not both.
📤 What you get back
A real row from a run on 2 October 2026, with the lists and the abstract cut short (the row carries 19 authors and 10 institutions):
{"ok": true,"openalexId": "https://openalex.org/W2101234009","doi": "https://doi.org/10.48550/arxiv.1201.0490","title": "Scikit-learn: Machine Learning in Python","authors": ["Fabián Pedregosa", "Gaël Varoquaux", "Alexandre Gramfort", "Michel, Vincent", "..."],"institutions": ["Commissariat à l'Énergie Atomique et aux Énergies Alternatives", "..."],"year": 2012,"publicationDate": "2012-01-02","type": "article","venue": "ORBi (University of Liège)","citations": 63573,"concepts": ["Python (programming language)", "Documentation", "Computer science", "MIT License", "Artificial intelligence"],"isOpenAccess": true,"oaUrl": "https://orbi.uliege.be/handle/2268/225787","abstract": "Scikit-learn is a Python module integrating a wide range of ... machine learning algorithms for medium-scale supervised and unsupervised problems. ...","url": "https://openalex.org/W2101234009"}
| Field | How to read it |
|---|---|
openalexId | OpenAlex's permanent id. Use it as your key, since not every work has a DOI. |
abstract | Rebuilt from OpenAlex's word-position index. null where there is no index to rebuild from. |
authors | Names as OpenAlex holds them, so a few arrive family name first, like Michel, Vincent. |
venue | OpenAlex's primary location for the work. Often the journal, but it can be a repository, as above, or null. |
citations | OpenAlex's cited-by count, and 0 when OpenAlex has no figure at all. |
institutions | The affiliations OpenAlex resolved for the authors, without repeats. Sometimes messy, because the matching is theirs. |
concepts | Up to five of OpenAlex's own topic tags, not keywords the authors chose. |
isOpenAccess, oaUrl | Whether a free copy is known, and where it is. oaUrl is null when none is known. |
type | article, preprint, book-chapter, dataset and so on, in OpenAlex's vocabulary. |
🧾 Reading the output
Two kinds of row land in your dataset, and ok tells them apart.
| Row | How to spot it | Billed |
|---|---|---|
| A work | ok: true and an openalexId | yes |
| A diagnostic | ok: false and an errorCode | no |
The overview table in the Apify Console shows the work columns only, so a diagnostic row looks blank there. Switch to the JSON or All fields view to read it.
| Code | What it means |
|---|---|
BAD_INPUT | No query was given. |
NO_RESULTS | The request worked and nothing came back. A malformed filter also lands here. |
RATE_LIMITED | OpenAlex asked for a slower pace than the run could keep. Try a smaller run. |
SERVER_ERROR | OpenAlex answered with a server error. Usually passes. |
BLOCKED | OpenAlex refused the request. Re-run it. |
NETWORK | OpenAlex was unreachable. Re-run it. |
💡 What people use it for
- Pulling a field's most-cited work with the abstracts attached, in one file, for a reading list.
- Feeding titles and abstracts to a model that has to summarise research it was not trained on.
- Finding which institutions keep appearing on a topic, using the
institutionsfield. - Filtering to open access with
is_oa:true, so every row has somewhere to read it for free. - Watching a topic on a schedule with
sortondate, so each run surfaces what is new.
🚧 What it does not do
- Keyword search over works only. No author, institution or funder lookup by entity.
- No full text. You get the abstract and a link, not the paper.
- Abstracts are missing where OpenAlex has no index for them, and that is common on older and paywalled records.
citationscannot tell you zero from unknown. Both read0.- A bad
filterlooks like an empty search, not like an error. - No end date field. Use the raw filter if you need an upper bound.
- Institutions and concepts are OpenAlex's matching, and it is not perfect. Treat them as hints, not as ground truth.
- All or nothing. A failure partway through discards what had been collected, so you get a diagnostic row rather than a partial set.
🧭 Which research scraper do you need?
| If you want | Use |
|---|---|
| arXiv preprints with full abstracts and PDF links | arXiv Scraper |
| Anything with a registered DOI, with its journal, publisher and ISSN | Crossref Scraper |
| Papers with readable abstracts, institutions and open-access links | This one |
| Patents rather than papers | Google Patents Search Scraper |
| Book records by keyword, title or ISBN | Books Scraper |
| Encyclopedia articles as plain text | Wikipedia Scraper |
| arXiv, OpenAlex and Wikipedia searches as tools for an AI agent | Research MCP Server |
❓ Questions people ask
Do I need an OpenAlex account or an API key?
No. OpenAlex is open and the actor reads it directly.
Why is the abstract null on some papers?
OpenAlex only keeps a word-position index for a work if the publisher made one available. No index, no abstract to rebuild.
Why is the venue a repository instead of a journal?
OpenAlex picks one primary location per work, and for a paper with a preprint or a repository copy it sometimes picks that. The DOI still points at the published version.
Can I search by author or institution?
Not as a mode. You can approximate it through the raw filter field, using OpenAlex's own filter
keys.
Can I call it from code or connect it to an AI assistant?
Yes. The API tab has ready-made code for
Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect
https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/openalex-scraper. Either way the run
happens on your Apify account at the same price.
Is scraping OpenAlex legal?
OpenAlex publishes this data openly under a public domain dedication and invites this kind of use. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID, the query and any filter you used. The
errorCode on the diagnostic row usually names the problem on its own.