Academic Papers Scraper — OpenAlex, 250M+ Works, No Login
Pricing
from $0.50 / 1,000 papers
Academic Papers Scraper — OpenAlex, 250M+ Works, No Login
Search 250M+ scholarly works from OpenAlex — by keyword, author, institution or journal; filter by year, citations, open access and work type. No login, no API key, no captcha. Normalized JSON: DOI, title, authors, venue, year, citations, OA PDF link, abstract.
Pricing
from $0.50 / 1,000 papers
Rating
0.0
(0)
Developer
Petro Pankov
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Share
Academic Papers Scraper — OpenAlex (250M+ works, no login)
Search the world's scholarly literature — journal articles, preprints, book chapters, datasets and reviews — and get clean, structured rows back. Built directly on OpenAlex, an open catalogue of over 250 million works.
No login. No API key. No captcha. No proxies needed.
Why this instead of a Google Scholar scraper
Google Scholar has no public API. Every scraper built on it is reverse-engineering an HTML surface that actively fights automation — which is why those Actors break, rate-limit, and need proxy budgets. OpenAlex publishes the same literature through a documented, open, versioned API with a stable schema. Nothing to fight, nothing to break.
That means: predictable runs, no proxy costs, no captcha failures, and results you can actually schedule against.
What you can do
- Keyword search across title, abstract and fulltext —
"large language models" - By author —
"Yoshua Bengio"(name match, no author IDs needed) - By institution —
"Massachusetts Institute of Technology","Max Planck" - By journal / conference —
"Nature","NeurIPS" - Filter by year range, minimum citations, open-access-only, and work type
- Sort by relevance, citation count, or newest first
Output
Each row is normalized:
| field | description |
|---|---|
id, openAlexUrl | OpenAlex work ID and URL |
doi | DOI (bare, no https://doi.org/ prefix) |
title | Work title |
authors, firstAuthor, authorCount | Full author list and convenience fields |
venue, publisher, issn | Journal / conference and publisher |
year, publishedAt | Publication year and ISO date |
type, language | Work type (article, preprint, …) and language |
citedByCount, referencedWorksCount, fwci | Citation metrics |
isOpenAccess, oaStatus, pdfUrl | Open-access status and direct free-PDF link |
landingPageUrl | Publisher landing page |
concepts, institutions, countries | Topics and affiliations |
abstract | Full abstract text (when Include abstract is on) |
Example
{"search": "retrieval augmented generation","fromYear": 2023,"minCitations": 10,"openAccessOnly": true,"sortBy": "citations","maxResults": 200,"includeAbstract": true}
Typical uses
- Literature reviews and systematic reviews
- Citation and research-trend analysis
- Tracking a lab's, author's or institution's output
- Building datasets of open-access PDFs for downstream NLP
- Competitive/landscape research in a scientific field
Notes
- Results are capped at today's date unless you set To year — OpenAlex contains a small number of records with future publication dates, and this keeps "newest first" meaningful.
- Requests go through OpenAlex's polite pool (a contact address is sent with each call), which is the usage pattern OpenAlex asks for.
- Abstracts are stored by OpenAlex as an inverted index for licensing reasons; this Actor reconstructs them into plain text for you.