Academic Papers Search API — OpenAlex, Crossref, arXiv, PubMed
Pricing
from $0.24 / 1,000 paper records
Academic Papers Search API — OpenAlex, Crossref, arXiv, PubMed
Search OpenAlex, Crossref, arXiv and PubMed, or look up DOIs, arXiv ids and PMIDs. One merged row per paper: title, abstract, venue, year, DOI, citations, open-access status and PDF link, authors with institutions, topics. Filter by year, type, open access, citations. No API key.
Pricing
from $0.24 / 1,000 paper records
Rating
0.0
(0)
Developer
Insight Solutions
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
9 hours ago
Last modified
Categories
Share
Academic Paper Search API: OpenAlex, Crossref, arXiv, PubMed
Search scholarly literature across four open sources and get one clean row per paper: title, abstract, year and date, type, venue and publisher, DOI and the other identifiers, citation count, open-access status with the PDF link when one exists, authors with their institutions, topics and keywords, language and the retraction flag.
| Source | What it is best at |
|---|---|
| OpenAlex (the default) | The broadest index: 250M+ works with abstracts, citation counts and percentiles, topics, open-access links and institutions |
| Crossref | The publishers' own metadata of record for a DOI: container, publisher, pages, licence, full-text links |
| arXiv | Preprints, with full abstracts, categories and a PDF on every paper |
| PubMed | Biomedicine, with MeSH terms, author keywords and structured abstracts |
Search by query with year, type, open-access and citation filters, or look up
DOIs, arXiv ids and PubMed ids directly. Results from several sources are
merged on DOI, so one paper is one row — and one charge — with sources
listing every source that had it.
No API key. No login. No browser. No e-mail address sent anywhere.
At a glance
Input — this is the Store prefill; paste it and run:
{ "queries": ["graphene battery"], "sources": ["openalex"], "yearFrom": null, "maxResults": 10,"sortBy": "relevance", "includeAbstract": true, "includeAuthors": true, "dedupe": true,"maxConcurrency": 3, "maxRunSecs": 240, "proxyConfiguration": { "useApifyProxy": true } }
That is one OpenAlex request and ten paper rows, in about two seconds.
Output — one paper row per work, the same columns from every source. The
fields you will use most are title, abstract, year, doi,
citationCount, isOpenAccess, pdfUrl, venue and authors (full list under
Output reference). Anything that could not be read — an unknown DOI, a search
with no match, a source that refused us — comes back as a free diagnostic row
(ok: false, errorType, error) instead of a charge.
Price — $0.40 per 1,000 papers (+ $0.001 per run); diagnostics and empty runs free; no API key, no browser, limited permissions, works over the Apify MCP server and with x402.
From code — client.actor("insight.solutions/academic-papers-api").call(run_input={…})
with apify-client, or
POST https://api.apify.com/v2/acts/insight.solutions~academic-papers-api/run-sync-get-dataset-items.
What you get
One row per paper. This is the first row of the prefill run, trimmed to the columns that carry something (two of its six authors and five of its keywords shown):
{"ok": true,"rowType": "paper","input": "graphene battery","source": "openalex","sourceUrl": "https://api.openalex.org/works/W2028157562","sources": ["openalex"],"title": "Large Reversible Li Storage of Graphene Nanosheet Families for Use in Rechargeable Lithium Ion Batteries","abstract": "The lithium storage properties of graphene nanosheet (GNS) materials as high capacity anode materials for rechargeable lithium secondary batteries (LIB) were in …","year": 2008,"publishedAt": "2008-07-24","type": "article","doi": "10.1021/nl800957b","doiUrl": "https://doi.org/10.1021/nl800957b","openalexId": "W2028157562","pmid": "18651781","landingUrl": "https://doi.org/10.1021/nl800957b","isOpenAccess": false,"oaStatus": "closed","venue": "Nano Letters","venueType": "journal","publisher": "American Chemical Society","issn": "1530-6984","volume": "8","issue": "8","pages": "2277-2282","citationCount": 2922,"referenceCount": 31,"fwci": 52.4896,"citationPercentile": 0.99948808,"isTop1Percent": true,"isRetracted": false,"language": "en","authors": [{ "name": "Eunjoo Yoo", "institutions": ["National Institute for Materials Science", "National Institute of Advanced Industrial Science and Technology"], "countries": ["JP"], "isGroup": false },{ "name": "Je‐Deok Kim", "institutions": ["National Institute for Materials Science", "National Institute of Advanced Industrial Science and Technology"], "countries": ["JP"], "isGroup": false }],"authorCount": 6,"firstAuthor": "Eunjoo Yoo","institutions": ["National Institute for Materials Science", "National Institute of Advanced Industrial Science and Technology"],"institutionCountries": ["JP"],"topics": ["Graphene research and applications", "Advancements in Battery Materials", "Advanced Battery Materials and Technologies"],"field": "Materials Science","subfield": "Materials Chemistry","domain": "Physical Sciences","keywords": ["Nanosheet", "Graphene", "Anode", "Lithium (medication)", "Graphite"],"query": "graphene battery","rank": 1,"updatedAt": "2026-09-25T08:18:29.962Z"}
An arXiv row fills arxivId, arxivVersion, arxivCategories, pdfUrl and
journalRef instead of the citation columns; a PubMed row fills pmid,
pmcid, meshTerms, keywords and publicationTypes, and its abstract keeps
the labelled sections (BACKGROUND: …, RESULTS: …), one per line.
Three dataset views ship with it: Papers, Open access (the PDF, landing page, OA status and licence columns) and Diagnostics.
Quick start
Search one source. The prefill above. Add "yearFrom": 2020,
"sortBy": "citations" or "openAccessOnly": true to narrow it.
A batch of DOIs, resolved through OpenAlex and, for anything OpenAlex lacks, Crossref:
{ "dois": ["10.1038/s41586-021-03819-2", "https://doi.org/10.1021/nl800957b", "doi:10.7551/mitpress/15517.003.0003"] }
An arXiv category, newest first — a query is optional; with none, the category is browsed on its own:
{ "arxivCategories": ["cs.CL"], "sortBy": "date", "maxResults": 50 }
PubMed, with abstracts and MeSH terms:
{ "queries": ["crispr base editing"], "sources": ["pubmed"], "yearFrom": 2023, "maxResults": 100 }
Every source at once, merged:
{ "queries": ["protein structure prediction"], "sources": ["openalex", "crossref", "arxiv", "pubmed"], "maxResults": 50 }
Input
| Field | Default | What it does | How each source applies it |
|---|---|---|---|
queries | [] | Free-text searches. One walk per query per source. | OpenAlex search=; Crossref query=; arXiv all:"…" (a query already in arXiv syntax, like ti:transformer AND au:vaswani, is passed through); PubMed term= |
sources | ["openalex"] | Where queries are searched. | Lookups always go to the source that owns the identifier. |
maxResults | 100 | Papers per query per source. 0 = no cap (the time budget still applies). | Paging: OpenAlex and Crossref cursors, arXiv start, PubMed retstart. Lookups are not capped. |
dois | [] | DOIs in any form (10.…, doi:…, https://doi.org/…). | OpenAlex works/doi:… first, Crossref works/… when OpenAlex has no record. |
arxivIds | [] | 2303.08774, 2303.08774v2, arXiv:…, arxiv.org/abs/… links, hep-th/9901001. | arXiv id_list, 50 per request. |
pmids | [] | PubMed ids, PMID: … or pubmed.ncbi.nlm.nih.gov links. | PubMed esummary + efetch, 50 per request. |
arxivCategories | [] | cs.CL, stat.ML, q-bio.NC… Adds arXiv to the sources. | ANDed onto each arXiv query; alone, a category browse. |
yearFrom / yearTo | — | Publication years, inclusive. | OpenAlex from_/to_publication_date; Crossref from-/until-pub-date; arXiv on the first-submission date, here; PubMed mindate/maxdate (datetype=pdat). |
types | [] (all) | article, preprint, review, book-chapter, book, dissertation, dataset, other. | OpenAlex and Crossref type: upstream (except other, and review on Crossref, which are checked here); arXiv is always preprint; PubMed from its publication types. |
openAccessOnly | false | Only papers free to read. | OpenAlex is_oa:true; arXiv always qualifies; PubMed when in PubMed Central; Crossref rows are dropped (it publishes no OA status) with one free skipped row. |
minCitations | — | Cited at least this often. | OpenAlex cited_by_count:>N-1; Crossref on its own count, here; arXiv and PubMed rows are kept with citationCount: null, because they publish no count. |
sortBy | relevance | relevance, citations or date. | Citations: OpenAlex and Crossref (arXiv and PubMed fall back to relevance). Date: all four, newest first. |
includeAbstract | true | Fill abstract. | OpenAlex rebuilt from its word index; Crossref when deposited; arXiv always; PubMed via one efetch per 50 papers (skipped when off). |
includeAuthors | true | Fill authors and firstAuthor. | Off: the list and the first author are dropped entirely, and OpenAlex is not even asked for them. |
dedupe | true | Merge one paper's records across sources. | See How duplicates are merged. Off: every source's row is pushed and charged. |
maxConcurrency | 3 | Walks in flight. | arXiv is always one request at a time, 3 s apart; PubMed at most 3 a second. |
maxRunSecs | 240 | Time budget, 30–3600 s. | Rows already returned are kept. |
proxyConfiguration | datacenter | Apify proxy. | All four sources also answered with no proxy at all. |
Filters apply to searches. An identifier lookup returns the record you asked for.
Output reference
Every row has the same 59 columns; a column that does not apply is null.
- Envelope:
ok,rowType(paperordiagnostic),input(the query or the identifier),source(the base record's source),sourceUrl(its API URL),sources(every source that had the paper),scrapedAt,error,errorType. - The paper:
title,abstract,year,publishedAt(YYYY-MM-DD,YYYY-MMorYYYY— as precise as the source),type(our vocabulary),typeRaw(the source's own word). - Identifiers:
doi,doiUrl,openalexId,arxivId,arxivVersion,pmid,pmcid. - Access:
landingUrl,pdfUrl,isOpenAccess,oaStatus,license. - Venue:
venue,venueType(journal,conference,repository,book-series,ebook-platform,other),publisher,issn,volume,issue,pages. - Impact:
citationCount,referenceCount,fwci,citationPercentile,isTop1Percent,isTop10Percent,isRetracted,language. - People:
authors[]({ name, institutions[], countries[], isGroup }),authorCount,firstAuthor,institutions[],institutionCountries[]. - Subject:
topics[](OpenAlex topics, Crossref subjects, arXiv categories or PubMed MeSH descriptors),field,subfield,domain(OpenAlex primary topic),keywords[],arxivCategories[],meshTerms[],publicationTypes[],journalRef. - Provenance:
query,rank(1-based position in the source's own result list, or in your identifier list),updatedAt(when the source last updated its record).
Diagnostic errorTypes, all free: invalid-id (not an identifier — no request
made), invalid-input (nothing usable to run), not-found, no-results,
bad-query (the source rejected the query, with its reason), rate-limited,
blocked, http, timeout, deadline, budget, upstream-format, skipped
(Crossref under openAccessOnly, or arXiv under a types filter without
preprint), and actor-error.
Which source for what
OpenAlex is the default and the one to start with. It indexes more than 250
million works across every discipline and carries, on one record, the abstract,
the citation count with field-normalised impact (fwci, percentile, top-1% and
top-10% flags), topics and keywords, the open-access verdict with the best PDF
link, and each author's institutions and countries. It applies every filter this
Actor offers upstream, except the catch-all other type.
Crossref is where publishers deposit their metadata of record. Use it next to OpenAlex when you want the publisher's own container title, page range, licence URL and full-text links, or on its own to resolve DOIs OpenAlex has not picked up yet. It has no open-access verdict and abstracts only when a publisher deposited one.
arXiv is preprints — physics, mathematics, computer science, statistics,
quantitative biology and finance, economics, electrical engineering. Every paper
has its full abstract and a PDF, and arxivCategories for precise topical
searches. There are no citation counts.
PubMed is biomedicine and life sciences. Its records carry MeSH descriptors, author keywords, publication types (so reviews and trials can be told apart) and structured abstracts; a paper in PubMed Central is marked open access.
How duplicates are merged
With dedupe on (the default) and more than one source in the run, the same work
found by several sources becomes one row. Records are matched on the lower-cased
DOI, then the arXiv id (OpenAlex's arXiv DOIs, 10.48550/arXiv.…, carry one),
then the PMID. The base record is chosen by source — OpenAlex, then Crossref,
then PubMed, then arXiv — never by which answered first, and the base keeps every
value it has; the other records only fill the base's empty columns. source names
the base, sources names all of them, and the row is charged once. A one-source
run streams its rows as they arrive and drops a paper a second query finds again.
Authors and privacy
Authors appear as printed on the paper: the display name, the names of the
institutions the source matched, and those institutions' country codes. Nothing
else. This Actor never outputs ORCID iDs, author ids, e-mail addresses,
corresponding-author flags, or the raw author-name and raw affiliation strings the
sources also carry (on real records those raw strings contain e-mail addresses).
Free text that we do pass through — an abstract, a publisher-deposited affiliation
line — has any e-mail address replaced with [e-mail removed]. A collective author
such as a study group or consortium is marked isGroup: true.
Set includeAuthors: false to drop the author list and firstAuthor entirely;
OpenAlex is then not even asked for them.
What you are never charged for
- Every diagnostic row —
invalid-id,not-found,no-results,bad-query,rate-limited,blocked,http,timeout,deadline,budget,upstream-format,skipped. - Duplicates. One paper found by three sources is one row and one charge; a paper a second query finds again is dropped, not billed.
- Papers your filters excluded, including Crossref rows dropped by
openAccessOnly. Every filter runs before the row is written. - The extra requests behind a row: PubMed's
efetch, the Crossref fallback for a DOI OpenAlex does not have, retries. - A run that returns nothing. A run whose input was usable but produced no paper — no match, an unknown identifier, filters that excluded everything, a source that refused us or was down, the time budget running out — finishes SUCCEEDED with zero results, with a status message that says so and points at the diagnostic rows, and bills nothing, start fee included. A run finishes FAILED only when there was nothing usable to attempt (no query and no valid identifier) or the Actor itself hit an error.
Pricing
| Event | FREE | BRONZE | SILVER | GOLD |
|---|---|---|---|---|
| Run started | $0.001 | $0.001 | $0.001 | $0.001 |
| Paper record | $0.0004 | $0.0004 | $0.00032 | $0.00024 |
$0.40 per 1,000 papers at the free tier, whichever source a paper came from. 500 papers cost 500 × $0.0004 + $0.001 = $0.201.
The run honours your maximum total charge (ACTOR_MAX_TOTAL_CHARGE_USD): when it
is reached the Actor stops fetching, leaves a free budget row, keeps everything
already delivered and finishes successfully.
Use it from an AI agent, or from code
One JSON object in, one flat array out. The Actor runs with limited permissions, uses pay-per-event pricing and never enters Standby, so it works over the Apify MCP server and with x402 agentic payments.
curl -X POST "https://api.apify.com/v2/acts/insight.solutions~academic-papers-api/run-sync-get-dataset-items?token=$APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"queries":["graphene battery"],"maxResults":25,"openAccessOnly":true}'
# pip install apify-clientfrom apify_client import ApifyClientclient = ApifyClient("<APIFY_TOKEN>")run = client.actor("insight.solutions/academic-papers-api").call(run_input={"queries": ["protein structure prediction"],"sources": ["openalex", "crossref", "arxiv", "pubmed"],"yearFrom": 2020,"sortBy": "citations","maxResults": 100,})for row in client.dataset(run["defaultDatasetId"]).iterate_items():if row["rowType"] != "paper":print("skipped:", row["errorType"], row["error"])continueprint(row["year"], row["citationCount"], row["title"], row["doi"], row["pdfUrl"], row["sources"], sep=" | ")
FAQ
Why do Crossref rows usually have no abstract?
Crossref only has an abstract when the publisher deposited one, and most do not:
1 of the 20 items in our captured Crossref search had one. OpenAlex rebuilds
abstracts for far more works, which is one reason it is the default — and with
dedupe on, a paper both sources have comes out with OpenAlex's abstract.
Why is arXiv slow?
arXiv's API rules ask for one request every three seconds, one at a time, and this
Actor keeps to that whatever maxConcurrency says. Each request returns up to 100
papers, so a 300-paper arXiv walk spends about six seconds waiting.
Why do citation counts differ between sources?
Each source counts over its own index. OpenAlex counted 47,662 citations of the
AlphaFold paper where Crossref counted 45,503 on the same day. A merged row keeps
the base record's count (OpenAlex's when OpenAlex has the paper); run with
dedupe: false to see both. arXiv and PubMed publish no counts at all.
How does dedupe choose the base record? By source, not by speed: OpenAlex, then Crossref, then PubMed, then arXiv. The base keeps all of its values and the others fill only its gaps, so the same input gives the same row every time.
Can I use my own OpenAlex or NCBI key, or the "polite pool"? No. This Actor sends no API key and no e-mail address; it runs in each source's anonymous pool and paces itself below the published limits.
Does minCitations drop arXiv and PubMed papers?
No — they have no count to compare, so they are kept with citationCount: null.
Use OpenAlex alone when a citation floor must be strict.
Limitations
- The anonymous pools have limits: OpenAlex allows 100,000 requests a day, NCBI
three a second, arXiv one every three seconds. A 429 is retried once, after a
pause and with a fresh proxy session; a second one becomes a free
rate-limitedrow. - PubMed's
esearchdoes not page past its 10,000th result; narrow the query or the years to reach further. - Crossref has no open-access status, so
openAccessOnlydrops Crossref rows, and it does not label review articles, sotypes: ["review"]matches no Crossref row. - A type filter of
otheris checked here rather than upstream, so such a walk may read many pages to find a few papers; it stops after ten pages in a row with nothing kept. - No full text: abstracts, metadata and links to the PDF only.
- Merging across sources holds a run's papers until every source has answered, up to 50,000 distinct papers per run.
- The upstream formats may change. The parsers are pinned against real responses
captured on 2026-09-28, and a response that no longer matches becomes a free
upstream-formatrow rather than a wrong one.
Sources, terms and attribution
- OpenAlex — metadata from OpenAlex (openalex.org), released under the CC0 public-domain dedication.
- Crossref — metadata from the Crossref REST API (api.crossref.org); Crossref makes its bibliographic metadata available without restriction, and abstracts remain the publishers' own.
- arXiv — thank you to arXiv for use of its open access interoperability; data is retrieved through the arXiv API under the arXiv API Terms of Use.
- PubMed — data from the National Library of Medicine's PubMed via the NCBI E-utilities, used under the NCBI usage policy; this Actor is not endorsed by NLM.
Our other Actors
Every Insight Solutions Actor is pay-per-result with no browser, no login and no API key, and every one of them returns free diagnostic rows instead of billing for failures. Prices are per 1,000 results.
Video, audio & social
- YouTube Transcript API — captions as timed segments, text, SRT or VTT, with language fallback and translation.
- YouTube Comments API — comments and replies with likes, pinned and hearted flags, newest or top sort.
- YouTube Channel API — a channel's videos, Shorts and live streams, plus YouTube search.
- Podcast Search, Episodes & Charts API — Apple Podcasts search, charts and full episode feeds.
- Bluesky Scraper — profiles, posts, followers and follows from the public AT Protocol API.
- Telegram Channel Scraper — posts, views and channel stats from public Telegram channels.
- Substack Scraper — posts with full free text, comments and publication profiles.
- Hacker News API — stories, comments, users, front page and a structured "Who is hiring?" parser from the official HN APIs.
- Discourse Forum API — topics, posts and categories from any Discourse community via its own JSON endpoints, usernames only.
News, documents & the web
- Google News Search, Topics & Real Article URLs — news search and topic feeds with the publisher's real URL decoded.
- Website to Markdown — Content Extractor for LLMs & RAG — any site as clean Markdown, text and heading-aware chunks.
- Internet Archive API — archive.org search, item metadata, files and reviews.
- Wayback Machine Toolkit — archived URL inventories, snapshots and text diffs between dates.
- Website Technology Detector — the tech stack behind any site, with the evidence for each detection.
- Domain Intelligence API — DNS, RDAP registration, TLS certificate and HTTP facts in one row per domain.
- SEO Page Audit — sitemap crawl with on-page checks, structured data and broken-link reports.
- Keyword Suggestions API — Google, YouTube, Bing, Amazon and eBay autocomplete with alphabet and question expansions.
- Website Contact Extractor — emails, phone numbers and social profiles from any list of websites.
- Web Search Results API — Bing and DuckDuckGo organic results with snippets, no key, no browser.
- Company Enrichment API — a domain in, a company profile out: firmographics, contacts, tech stack, DNS and hiring signal.
- Company Dossier API — one company in, twelve sections out: profile, tech, contacts, DNS, open roles, news, SEC filings, federal awards, recalls, YC batch and apps.
- Press Releases API — GlobeNewswire and PR Newswire releases plus any newsroom feed, by keyword, company, ticker or subject.
- Federal Register API — rules, proposed rules, notices and the Public Inspection desk with dockets, comment deadlines and CFR references.
- RSS & Atom Feed Monitor — any RSS, Atom or JSON feed (or an OPML file) in, only the new items out, with keyword filters and a webhook.
- Website Change Monitor — watch any pages, diff the text between runs, get change rows with added/removed lines, keyword alerts and a webhook.
- Wikipedia & Wikidata API — article text, search, daily pageviews and Wikidata entity facts, any language edition.
Business, finance & jobs
- Congress & Insider Trades API — STOCK Act periodic transaction reports and SEC Form 4 insider trades in one schema.
- Federal Contracts, Grants & Lobbying API — SAM.gov opportunities, USAspending awards, Grants.gov notices and Senate lobbying filings in one schema.
- SEC EDGAR API — filings, XBRL financials and full-text search by ticker or CIK.
- Clinical Trials & FDA API — ClinicalTrials.gov studies plus openFDA recalls, labels, approvals, 510(k)s and adverse-event reports.
- Product & Vehicle Recalls API — CPSC, NHTSA, FDA and USDA recalls, vehicle complaints and ratings, plus a VIN decoder.
- Y Combinator Companies, Batches & Founders — the YC directory with founders and social links, filterable by batch, industry and hiring status.
- Career Site Jobs API — jobs straight from Greenhouse, Lever, Ashby, Workable and 10+ other ATS career sites.
- New Job Postings Monitor — new, closed and changed postings on the career sites you watch.
- Hiring Signals API — Open Roles & Hiring Surge by Company — one row per company per run: open roles, what opened and closed, department and seniority breakdowns, and a hiring-surge flag.
- Remote Jobs API — RemoteOK, Remotive, We Work Remotely, Himalayas, Jobicy and more in one schema, deduplicated.
- Shopify Products API — any Shopify store's catalogue, variants, prices and stock signals.
- Shopify Store Monitor — price drops, sales, restocks, sell-outs and new products on any Shopify store, one row per change.
- Public Tenders API — EU TED, UK Find a Tender and Contracts Finder notices by keyword, CPV code, country, stage and deadline.
- Nonprofit & IRS 990 Lookup API — search US nonprofits and get EIN, NTEE code and multi-year Form 990 financials.
- OpenStreetMap Places API — businesses and points of interest by category and area from OpenStreetMap: name, address, coordinates, website, phone, opening hours.
Apps & games
- App Store & Google Play Reviews API — reviews from both stores with ratings, versions and developer replies.
- App Store Top Charts & App Search API — Apple top charts by country and genre, plus app search and details.
- App Store Keyword Rank Tracker — where any app ranks for any keyword on the App Store and Google Play, with rank changes and ASO suggestions.
- Steam Reviews API — Steam reviews with playtime, helpfulness and game details.
- Steam Game Data API — prices, tags, review scores, live player counts and top charts.