Academic Papers Search API — OpenAlex, Crossref, arXiv, PubMed avatar

Academic Papers Search API — OpenAlex, Crossref, arXiv, PubMed

Pricing

from $0.24 / 1,000 paper records

Go to Apify Store
Academic Papers Search API — OpenAlex, Crossref, arXiv, PubMed

Academic Papers Search API — OpenAlex, Crossref, arXiv, PubMed

Search OpenAlex, Crossref, arXiv and PubMed, or look up DOIs, arXiv ids and PMIDs. One merged row per paper: title, abstract, venue, year, DOI, citations, open-access status and PDF link, authors with institutions, topics. Filter by year, type, open access, citations. No API key.

Pricing

from $0.24 / 1,000 paper records

Rating

0.0

(0)

Developer

Insight Solutions

Insight Solutions

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

9 hours ago

Last modified

Share

Academic Paper Search API: OpenAlex, Crossref, arXiv, PubMed

Search scholarly literature across four open sources and get one clean row per paper: title, abstract, year and date, type, venue and publisher, DOI and the other identifiers, citation count, open-access status with the PDF link when one exists, authors with their institutions, topics and keywords, language and the retraction flag.

SourceWhat it is best at
OpenAlex (the default)The broadest index: 250M+ works with abstracts, citation counts and percentiles, topics, open-access links and institutions
CrossrefThe publishers' own metadata of record for a DOI: container, publisher, pages, licence, full-text links
arXivPreprints, with full abstracts, categories and a PDF on every paper
PubMedBiomedicine, with MeSH terms, author keywords and structured abstracts

Search by query with year, type, open-access and citation filters, or look up DOIs, arXiv ids and PubMed ids directly. Results from several sources are merged on DOI, so one paper is one row — and one charge — with sources listing every source that had it.

No API key. No login. No browser. No e-mail address sent anywhere.

At a glance

Input — this is the Store prefill; paste it and run:

{ "queries": ["graphene battery"], "sources": ["openalex"], "yearFrom": null, "maxResults": 10,
"sortBy": "relevance", "includeAbstract": true, "includeAuthors": true, "dedupe": true,
"maxConcurrency": 3, "maxRunSecs": 240, "proxyConfiguration": { "useApifyProxy": true } }

That is one OpenAlex request and ten paper rows, in about two seconds.

Output — one paper row per work, the same columns from every source. The fields you will use most are title, abstract, year, doi, citationCount, isOpenAccess, pdfUrl, venue and authors (full list under Output reference). Anything that could not be read — an unknown DOI, a search with no match, a source that refused us — comes back as a free diagnostic row (ok: false, errorType, error) instead of a charge.

Price — $0.40 per 1,000 papers (+ $0.001 per run); diagnostics and empty runs free; no API key, no browser, limited permissions, works over the Apify MCP server and with x402.

From code — client.actor("insight.solutions/academic-papers-api").call(run_input={…}) with apify-client, or POST https://api.apify.com/v2/acts/insight.solutions~academic-papers-api/run-sync-get-dataset-items.


What you get

One row per paper. This is the first row of the prefill run, trimmed to the columns that carry something (two of its six authors and five of its keywords shown):

{
"ok": true,
"rowType": "paper",
"input": "graphene battery",
"source": "openalex",
"sourceUrl": "https://api.openalex.org/works/W2028157562",
"sources": ["openalex"],
"title": "Large Reversible Li Storage of Graphene Nanosheet Families for Use in Rechargeable Lithium Ion Batteries",
"abstract": "The lithium storage properties of graphene nanosheet (GNS) materials as high capacity anode materials for rechargeable lithium secondary batteries (LIB) were in …",
"year": 2008,
"publishedAt": "2008-07-24",
"type": "article",
"doi": "10.1021/nl800957b",
"doiUrl": "https://doi.org/10.1021/nl800957b",
"openalexId": "W2028157562",
"pmid": "18651781",
"landingUrl": "https://doi.org/10.1021/nl800957b",
"isOpenAccess": false,
"oaStatus": "closed",
"venue": "Nano Letters",
"venueType": "journal",
"publisher": "American Chemical Society",
"issn": "1530-6984",
"volume": "8",
"issue": "8",
"pages": "2277-2282",
"citationCount": 2922,
"referenceCount": 31,
"fwci": 52.4896,
"citationPercentile": 0.99948808,
"isTop1Percent": true,
"isRetracted": false,
"language": "en",
"authors": [
{ "name": "Eunjoo Yoo", "institutions": ["National Institute for Materials Science", "National Institute of Advanced Industrial Science and Technology"], "countries": ["JP"], "isGroup": false },
{ "name": "Je‐Deok Kim", "institutions": ["National Institute for Materials Science", "National Institute of Advanced Industrial Science and Technology"], "countries": ["JP"], "isGroup": false }
],
"authorCount": 6,
"firstAuthor": "Eunjoo Yoo",
"institutions": ["National Institute for Materials Science", "National Institute of Advanced Industrial Science and Technology"],
"institutionCountries": ["JP"],
"topics": ["Graphene research and applications", "Advancements in Battery Materials", "Advanced Battery Materials and Technologies"],
"field": "Materials Science",
"subfield": "Materials Chemistry",
"domain": "Physical Sciences",
"keywords": ["Nanosheet", "Graphene", "Anode", "Lithium (medication)", "Graphite"],
"query": "graphene battery",
"rank": 1,
"updatedAt": "2026-09-25T08:18:29.962Z"
}

An arXiv row fills arxivId, arxivVersion, arxivCategories, pdfUrl and journalRef instead of the citation columns; a PubMed row fills pmid, pmcid, meshTerms, keywords and publicationTypes, and its abstract keeps the labelled sections (BACKGROUND: …, RESULTS: …), one per line.

Three dataset views ship with it: Papers, Open access (the PDF, landing page, OA status and licence columns) and Diagnostics.


Quick start

Search one source. The prefill above. Add "yearFrom": 2020, "sortBy": "citations" or "openAccessOnly": true to narrow it.

A batch of DOIs, resolved through OpenAlex and, for anything OpenAlex lacks, Crossref:

{ "dois": ["10.1038/s41586-021-03819-2", "https://doi.org/10.1021/nl800957b", "doi:10.7551/mitpress/15517.003.0003"] }

An arXiv category, newest first — a query is optional; with none, the category is browsed on its own:

{ "arxivCategories": ["cs.CL"], "sortBy": "date", "maxResults": 50 }

PubMed, with abstracts and MeSH terms:

{ "queries": ["crispr base editing"], "sources": ["pubmed"], "yearFrom": 2023, "maxResults": 100 }

Every source at once, merged:

{ "queries": ["protein structure prediction"], "sources": ["openalex", "crossref", "arxiv", "pubmed"], "maxResults": 50 }

Input

FieldDefaultWhat it doesHow each source applies it
queries[]Free-text searches. One walk per query per source.OpenAlex search=; Crossref query=; arXiv all:"…" (a query already in arXiv syntax, like ti:transformer AND au:vaswani, is passed through); PubMed term=
sources["openalex"]Where queries are searched.Lookups always go to the source that owns the identifier.
maxResults100Papers per query per source. 0 = no cap (the time budget still applies).Paging: OpenAlex and Crossref cursors, arXiv start, PubMed retstart. Lookups are not capped.
dois[]DOIs in any form (10.…, doi:…, https://doi.org/…).OpenAlex works/doi:… first, Crossref works/… when OpenAlex has no record.
arxivIds[]2303.08774, 2303.08774v2, arXiv:…, arxiv.org/abs/… links, hep-th/9901001.arXiv id_list, 50 per request.
pmids[]PubMed ids, PMID: … or pubmed.ncbi.nlm.nih.gov links.PubMed esummary + efetch, 50 per request.
arxivCategories[]cs.CL, stat.ML, q-bio.NC… Adds arXiv to the sources.ANDed onto each arXiv query; alone, a category browse.
yearFrom / yearTo—Publication years, inclusive.OpenAlex from_/to_publication_date; Crossref from-/until-pub-date; arXiv on the first-submission date, here; PubMed mindate/maxdate (datetype=pdat).
types[] (all)article, preprint, review, book-chapter, book, dissertation, dataset, other.OpenAlex and Crossref type: upstream (except other, and review on Crossref, which are checked here); arXiv is always preprint; PubMed from its publication types.
openAccessOnlyfalseOnly papers free to read.OpenAlex is_oa:true; arXiv always qualifies; PubMed when in PubMed Central; Crossref rows are dropped (it publishes no OA status) with one free skipped row.
minCitations—Cited at least this often.OpenAlex cited_by_count:>N-1; Crossref on its own count, here; arXiv and PubMed rows are kept with citationCount: null, because they publish no count.
sortByrelevancerelevance, citations or date.Citations: OpenAlex and Crossref (arXiv and PubMed fall back to relevance). Date: all four, newest first.
includeAbstracttrueFill abstract.OpenAlex rebuilt from its word index; Crossref when deposited; arXiv always; PubMed via one efetch per 50 papers (skipped when off).
includeAuthorstrueFill authors and firstAuthor.Off: the list and the first author are dropped entirely, and OpenAlex is not even asked for them.
dedupetrueMerge one paper's records across sources.See How duplicates are merged. Off: every source's row is pushed and charged.
maxConcurrency3Walks in flight.arXiv is always one request at a time, 3 s apart; PubMed at most 3 a second.
maxRunSecs240Time budget, 30–3600 s.Rows already returned are kept.
proxyConfigurationdatacenterApify proxy.All four sources also answered with no proxy at all.

Filters apply to searches. An identifier lookup returns the record you asked for.


Output reference

Every row has the same 59 columns; a column that does not apply is null.

  • Envelope: ok, rowType (paper or diagnostic), input (the query or the identifier), source (the base record's source), sourceUrl (its API URL), sources (every source that had the paper), scrapedAt, error, errorType.
  • The paper: title, abstract, year, publishedAt (YYYY-MM-DD, YYYY-MM or YYYY — as precise as the source), type (our vocabulary), typeRaw (the source's own word).
  • Identifiers: doi, doiUrl, openalexId, arxivId, arxivVersion, pmid, pmcid.
  • Access: landingUrl, pdfUrl, isOpenAccess, oaStatus, license.
  • Venue: venue, venueType (journal, conference, repository, book-series, ebook-platform, other), publisher, issn, volume, issue, pages.
  • Impact: citationCount, referenceCount, fwci, citationPercentile, isTop1Percent, isTop10Percent, isRetracted, language.
  • People: authors[] ({ name, institutions[], countries[], isGroup }), authorCount, firstAuthor, institutions[], institutionCountries[].
  • Subject: topics[] (OpenAlex topics, Crossref subjects, arXiv categories or PubMed MeSH descriptors), field, subfield, domain (OpenAlex primary topic), keywords[], arxivCategories[], meshTerms[], publicationTypes[], journalRef.
  • Provenance: query, rank (1-based position in the source's own result list, or in your identifier list), updatedAt (when the source last updated its record).

Diagnostic errorTypes, all free: invalid-id (not an identifier — no request made), invalid-input (nothing usable to run), not-found, no-results, bad-query (the source rejected the query, with its reason), rate-limited, blocked, http, timeout, deadline, budget, upstream-format, skipped (Crossref under openAccessOnly, or arXiv under a types filter without preprint), and actor-error.


Which source for what

OpenAlex is the default and the one to start with. It indexes more than 250 million works across every discipline and carries, on one record, the abstract, the citation count with field-normalised impact (fwci, percentile, top-1% and top-10% flags), topics and keywords, the open-access verdict with the best PDF link, and each author's institutions and countries. It applies every filter this Actor offers upstream, except the catch-all other type.

Crossref is where publishers deposit their metadata of record. Use it next to OpenAlex when you want the publisher's own container title, page range, licence URL and full-text links, or on its own to resolve DOIs OpenAlex has not picked up yet. It has no open-access verdict and abstracts only when a publisher deposited one.

arXiv is preprints — physics, mathematics, computer science, statistics, quantitative biology and finance, economics, electrical engineering. Every paper has its full abstract and a PDF, and arxivCategories for precise topical searches. There are no citation counts.

PubMed is biomedicine and life sciences. Its records carry MeSH descriptors, author keywords, publication types (so reviews and trials can be told apart) and structured abstracts; a paper in PubMed Central is marked open access.

How duplicates are merged

With dedupe on (the default) and more than one source in the run, the same work found by several sources becomes one row. Records are matched on the lower-cased DOI, then the arXiv id (OpenAlex's arXiv DOIs, 10.48550/arXiv.…, carry one), then the PMID. The base record is chosen by source — OpenAlex, then Crossref, then PubMed, then arXiv — never by which answered first, and the base keeps every value it has; the other records only fill the base's empty columns. source names the base, sources names all of them, and the row is charged once. A one-source run streams its rows as they arrive and drops a paper a second query finds again.

Authors and privacy

Authors appear as printed on the paper: the display name, the names of the institutions the source matched, and those institutions' country codes. Nothing else. This Actor never outputs ORCID iDs, author ids, e-mail addresses, corresponding-author flags, or the raw author-name and raw affiliation strings the sources also carry (on real records those raw strings contain e-mail addresses). Free text that we do pass through — an abstract, a publisher-deposited affiliation line — has any e-mail address replaced with [e-mail removed]. A collective author such as a study group or consortium is marked isGroup: true.

Set includeAuthors: false to drop the author list and firstAuthor entirely; OpenAlex is then not even asked for them.

What you are never charged for

  • Every diagnostic row — invalid-id, not-found, no-results, bad-query, rate-limited, blocked, http, timeout, deadline, budget, upstream-format, skipped.
  • Duplicates. One paper found by three sources is one row and one charge; a paper a second query finds again is dropped, not billed.
  • Papers your filters excluded, including Crossref rows dropped by openAccessOnly. Every filter runs before the row is written.
  • The extra requests behind a row: PubMed's efetch, the Crossref fallback for a DOI OpenAlex does not have, retries.
  • A run that returns nothing. A run whose input was usable but produced no paper — no match, an unknown identifier, filters that excluded everything, a source that refused us or was down, the time budget running out — finishes SUCCEEDED with zero results, with a status message that says so and points at the diagnostic rows, and bills nothing, start fee included. A run finishes FAILED only when there was nothing usable to attempt (no query and no valid identifier) or the Actor itself hit an error.

Pricing

EventFREEBRONZESILVERGOLD
Run started$0.001$0.001$0.001$0.001
Paper record$0.0004$0.0004$0.00032$0.00024

$0.40 per 1,000 papers at the free tier, whichever source a paper came from. 500 papers cost 500 × $0.0004 + $0.001 = $0.201.

The run honours your maximum total charge (ACTOR_MAX_TOTAL_CHARGE_USD): when it is reached the Actor stops fetching, leaves a free budget row, keeps everything already delivered and finishes successfully.


Use it from an AI agent, or from code

One JSON object in, one flat array out. The Actor runs with limited permissions, uses pay-per-event pricing and never enters Standby, so it works over the Apify MCP server and with x402 agentic payments.

curl -X POST "https://api.apify.com/v2/acts/insight.solutions~academic-papers-api/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"queries":["graphene battery"],"maxResults":25,"openAccessOnly":true}'
# pip install apify-client
from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("insight.solutions/academic-papers-api").call(run_input={
"queries": ["protein structure prediction"],
"sources": ["openalex", "crossref", "arxiv", "pubmed"],
"yearFrom": 2020,
"sortBy": "citations",
"maxResults": 100,
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
if row["rowType"] != "paper":
print("skipped:", row["errorType"], row["error"])
continue
print(row["year"], row["citationCount"], row["title"], row["doi"], row["pdfUrl"], row["sources"], sep=" | ")

FAQ

Why do Crossref rows usually have no abstract? Crossref only has an abstract when the publisher deposited one, and most do not: 1 of the 20 items in our captured Crossref search had one. OpenAlex rebuilds abstracts for far more works, which is one reason it is the default — and with dedupe on, a paper both sources have comes out with OpenAlex's abstract.

Why is arXiv slow? arXiv's API rules ask for one request every three seconds, one at a time, and this Actor keeps to that whatever maxConcurrency says. Each request returns up to 100 papers, so a 300-paper arXiv walk spends about six seconds waiting.

Why do citation counts differ between sources? Each source counts over its own index. OpenAlex counted 47,662 citations of the AlphaFold paper where Crossref counted 45,503 on the same day. A merged row keeps the base record's count (OpenAlex's when OpenAlex has the paper); run with dedupe: false to see both. arXiv and PubMed publish no counts at all.

How does dedupe choose the base record? By source, not by speed: OpenAlex, then Crossref, then PubMed, then arXiv. The base keeps all of its values and the others fill only its gaps, so the same input gives the same row every time.

Can I use my own OpenAlex or NCBI key, or the "polite pool"? No. This Actor sends no API key and no e-mail address; it runs in each source's anonymous pool and paces itself below the published limits.

Does minCitations drop arXiv and PubMed papers? No — they have no count to compare, so they are kept with citationCount: null. Use OpenAlex alone when a citation floor must be strict.

Limitations

  • The anonymous pools have limits: OpenAlex allows 100,000 requests a day, NCBI three a second, arXiv one every three seconds. A 429 is retried once, after a pause and with a fresh proxy session; a second one becomes a free rate-limited row.
  • PubMed's esearch does not page past its 10,000th result; narrow the query or the years to reach further.
  • Crossref has no open-access status, so openAccessOnly drops Crossref rows, and it does not label review articles, so types: ["review"] matches no Crossref row.
  • A type filter of other is checked here rather than upstream, so such a walk may read many pages to find a few papers; it stops after ten pages in a row with nothing kept.
  • No full text: abstracts, metadata and links to the PDF only.
  • Merging across sources holds a run's papers until every source has answered, up to 50,000 distinct papers per run.
  • The upstream formats may change. The parsers are pinned against real responses captured on 2026-09-28, and a response that no longer matches becomes a free upstream-format row rather than a wrong one.

Sources, terms and attribution

  • OpenAlex — metadata from OpenAlex (openalex.org), released under the CC0 public-domain dedication.
  • Crossref — metadata from the Crossref REST API (api.crossref.org); Crossref makes its bibliographic metadata available without restriction, and abstracts remain the publishers' own.
  • arXiv — thank you to arXiv for use of its open access interoperability; data is retrieved through the arXiv API under the arXiv API Terms of Use.
  • PubMed — data from the National Library of Medicine's PubMed via the NCBI E-utilities, used under the NCBI usage policy; this Actor is not endorsed by NLM.

Our other Actors

Every Insight Solutions Actor is pay-per-result with no browser, no login and no API key, and every one of them returns free diagnostic rows instead of billing for failures. Prices are per 1,000 results.

Video, audio & social

News, documents & the web

Business, finance & jobs

Apps & games