# Academic Papers Search API — OpenAlex, Crossref, arXiv, PubMed (`insight.solutions/academic-papers-api`) Actor

Search OpenAlex, Crossref, arXiv and PubMed, or look up DOIs, arXiv ids and PMIDs. One merged row per paper: title, abstract, venue, year, DOI, citations, open-access status and PDF link, authors with institutions, topics. Filter by year, type, open access, citations. No API key.

- **URL**: https://apify.com/insight.solutions/academic-papers-api.md
- **Developed by:** [Insight Solutions](https://apify.com/insight.solutions) (community)
- **Categories:** AI, Developer tools, Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.24 / 1,000 paper records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Academic Paper Search API: OpenAlex, Crossref, arXiv, PubMed

Search scholarly literature across **four open sources** and get **one clean
row per paper**: title, abstract, year and date, type, venue and publisher, DOI
and the other identifiers, citation count, open-access status with the PDF link
when one exists, authors with their institutions, topics and keywords, language
and the retraction flag.

| Source | What it is best at |
|---|---|
| **OpenAlex** (the default) | The broadest index: 250M+ works with abstracts, citation counts and percentiles, topics, open-access links and institutions |
| **Crossref** | The publishers' own metadata of record for a DOI: container, publisher, pages, licence, full-text links |
| **arXiv** | Preprints, with full abstracts, categories and a PDF on every paper |
| **PubMed** | Biomedicine, with MeSH terms, author keywords and structured abstracts |

Search by query with year, type, open-access and citation filters, or look up
DOIs, arXiv ids and PubMed ids directly. Results from several sources are
**merged on DOI**, so one paper is one row — and one charge — with `sources`
listing every source that had it.

No API key. No login. No browser. No e-mail address sent anywhere.

### At a glance

**Input** — this is the Store prefill; paste it and run:

```json
{ "queries": ["graphene battery"], "sources": ["openalex"], "yearFrom": null, "maxResults": 10,
  "sortBy": "relevance", "includeAbstract": true, "includeAuthors": true, "dedupe": true,
  "maxConcurrency": 3, "maxRunSecs": 240, "proxyConfiguration": { "useApifyProxy": true } }
```

That is one OpenAlex request and ten `paper` rows, in about two seconds.

**Output** — one `paper` row per work, the same columns from every source. The
fields you will use most are `title`, `abstract`, `year`, `doi`,
`citationCount`, `isOpenAccess`, `pdfUrl`, `venue` and `authors` (full list under
*Output reference*). Anything that could not be read — an unknown DOI, a search
with no match, a source that refused us — comes back as a free diagnostic row
(`ok: false`, `errorType`, `error`) instead of a charge.

**Price** — $0.40 per 1,000 papers (+ $0.001 per run); diagnostics and empty
runs free; no API key, no browser, limited permissions, works over the Apify MCP
server and with x402.

**From code** — `client.actor("insight.solutions/academic-papers-api").call(run_input={…})`
with `apify-client`, or
`POST https://api.apify.com/v2/acts/insight.solutions~academic-papers-api/run-sync-get-dataset-items`.

***

### What you get

One row per paper. This is the first row of the prefill run, trimmed to the
columns that carry something (two of its six authors and five of its keywords
shown):

```json
{
  "ok": true,
  "rowType": "paper",
  "input": "graphene battery",
  "source": "openalex",
  "sourceUrl": "https://api.openalex.org/works/W2028157562",
  "sources": ["openalex"],
  "title": "Large Reversible Li Storage of Graphene Nanosheet Families for Use in Rechargeable Lithium Ion Batteries",
  "abstract": "The lithium storage properties of graphene nanosheet (GNS) materials as high capacity anode materials for rechargeable lithium secondary batteries (LIB) were in …",
  "year": 2008,
  "publishedAt": "2008-07-24",
  "type": "article",
  "doi": "10.1021/nl800957b",
  "doiUrl": "https://doi.org/10.1021/nl800957b",
  "openalexId": "W2028157562",
  "pmid": "18651781",
  "landingUrl": "https://doi.org/10.1021/nl800957b",
  "isOpenAccess": false,
  "oaStatus": "closed",
  "venue": "Nano Letters",
  "venueType": "journal",
  "publisher": "American Chemical Society",
  "issn": "1530-6984",
  "volume": "8",
  "issue": "8",
  "pages": "2277-2282",
  "citationCount": 2922,
  "referenceCount": 31,
  "fwci": 52.4896,
  "citationPercentile": 0.99948808,
  "isTop1Percent": true,
  "isRetracted": false,
  "language": "en",
  "authors": [
    { "name": "Eunjoo Yoo", "institutions": ["National Institute for Materials Science", "National Institute of Advanced Industrial Science and Technology"], "countries": ["JP"], "isGroup": false },
    { "name": "Je‐Deok Kim", "institutions": ["National Institute for Materials Science", "National Institute of Advanced Industrial Science and Technology"], "countries": ["JP"], "isGroup": false }
  ],
  "authorCount": 6,
  "firstAuthor": "Eunjoo Yoo",
  "institutions": ["National Institute for Materials Science", "National Institute of Advanced Industrial Science and Technology"],
  "institutionCountries": ["JP"],
  "topics": ["Graphene research and applications", "Advancements in Battery Materials", "Advanced Battery Materials and Technologies"],
  "field": "Materials Science",
  "subfield": "Materials Chemistry",
  "domain": "Physical Sciences",
  "keywords": ["Nanosheet", "Graphene", "Anode", "Lithium (medication)", "Graphite"],
  "query": "graphene battery",
  "rank": 1,
  "updatedAt": "2026-09-25T08:18:29.962Z"
}
```

An arXiv row fills `arxivId`, `arxivVersion`, `arxivCategories`, `pdfUrl` and
`journalRef` instead of the citation columns; a PubMed row fills `pmid`,
`pmcid`, `meshTerms`, `keywords` and `publicationTypes`, and its abstract keeps
the labelled sections (`BACKGROUND: …`, `RESULTS: …`), one per line.

Three dataset views ship with it: **Papers**, **Open access** (the PDF, landing
page, OA status and licence columns) and **Diagnostics**.

***

### Quick start

**Search one source.** The prefill above. Add `"yearFrom": 2020`,
`"sortBy": "citations"` or `"openAccessOnly": true` to narrow it.

**A batch of DOIs**, resolved through OpenAlex and, for anything OpenAlex lacks,
Crossref:

```json
{ "dois": ["10.1038/s41586-021-03819-2", "https://doi.org/10.1021/nl800957b", "doi:10.7551/mitpress/15517.003.0003"] }
```

**An arXiv category**, newest first — a query is optional; with none, the
category is browsed on its own:

```json
{ "arxivCategories": ["cs.CL"], "sortBy": "date", "maxResults": 50 }
```

**PubMed, with abstracts and MeSH terms**:

```json
{ "queries": ["crispr base editing"], "sources": ["pubmed"], "yearFrom": 2023, "maxResults": 100 }
```

**Every source at once, merged**:

```json
{ "queries": ["protein structure prediction"], "sources": ["openalex", "crossref", "arxiv", "pubmed"], "maxResults": 50 }
```

***

### Input

| Field | Default | What it does | How each source applies it |
|---|---|---|---|
| `queries` | `[]` | Free-text searches. One walk per query per source. | OpenAlex `search=`; Crossref `query=`; arXiv `all:"…"` (a query already in arXiv syntax, like `ti:transformer AND au:vaswani`, is passed through); PubMed `term=` |
| `sources` | `["openalex"]` | Where `queries` are searched. | Lookups always go to the source that owns the identifier. |
| `maxResults` | `100` | Papers per query per source. `0` = no cap (the time budget still applies). | Paging: OpenAlex and Crossref cursors, arXiv `start`, PubMed `retstart`. Lookups are not capped. |
| `dois` | `[]` | DOIs in any form (`10.…`, `doi:…`, `https://doi.org/…`). | OpenAlex `works/doi:…` first, Crossref `works/…` when OpenAlex has no record. |
| `arxivIds` | `[]` | `2303.08774`, `2303.08774v2`, `arXiv:…`, `arxiv.org/abs/…` links, `hep-th/9901001`. | arXiv `id_list`, 50 per request. |
| `pmids` | `[]` | PubMed ids, `PMID: …` or pubmed.ncbi.nlm.nih.gov links. | PubMed `esummary` + `efetch`, 50 per request. |
| `arxivCategories` | `[]` | `cs.CL`, `stat.ML`, `q-bio.NC`… Adds arXiv to the sources. | ANDed onto each arXiv query; alone, a category browse. |
| `yearFrom` / `yearTo` | — | Publication years, inclusive. | OpenAlex `from_/to_publication_date`; Crossref `from-/until-pub-date`; arXiv on the first-submission date, here; PubMed `mindate`/`maxdate` (`datetype=pdat`). |
| `types` | `[]` (all) | `article`, `preprint`, `review`, `book-chapter`, `book`, `dissertation`, `dataset`, `other`. | OpenAlex and Crossref `type:` upstream (except `other`, and `review` on Crossref, which are checked here); arXiv is always `preprint`; PubMed from its publication types. |
| `openAccessOnly` | `false` | Only papers free to read. | OpenAlex `is_oa:true`; arXiv always qualifies; PubMed when in PubMed Central; **Crossref rows are dropped** (it publishes no OA status) with one free `skipped` row. |
| `minCitations` | — | Cited at least this often. | OpenAlex `cited_by_count:>N-1`; Crossref on its own count, here; **arXiv and PubMed rows are kept** with `citationCount: null`, because they publish no count. |
| `sortBy` | `relevance` | `relevance`, `citations` or `date`. | Citations: OpenAlex and Crossref (arXiv and PubMed fall back to relevance). Date: all four, newest first. |
| `includeAbstract` | `true` | Fill `abstract`. | OpenAlex rebuilt from its word index; Crossref when deposited; arXiv always; PubMed via one `efetch` per 50 papers (skipped when off). |
| `includeAuthors` | `true` | Fill `authors` and `firstAuthor`. | Off: the list and the first author are dropped entirely, and OpenAlex is not even asked for them. |
| `dedupe` | `true` | Merge one paper's records across sources. | See *How duplicates are merged*. Off: every source's row is pushed and charged. |
| `maxConcurrency` | `3` | Walks in flight. | arXiv is always one request at a time, 3 s apart; PubMed at most 3 a second. |
| `maxRunSecs` | `240` | Time budget, 30–3600 s. | Rows already returned are kept. |
| `proxyConfiguration` | datacenter | Apify proxy. | All four sources also answered with no proxy at all. |

Filters apply to searches. An identifier lookup returns the record you asked for.

***

### Output reference

Every row has the same 59 columns; a column that does not apply is `null`.

- **Envelope:** `ok`, `rowType` (`paper` or `diagnostic`), `input` (the query or
  the identifier), `source` (the base record's source), `sourceUrl` (its API
  URL), `sources` (every source that had the paper), `scrapedAt`, `error`,
  `errorType`.
- **The paper:** `title`, `abstract`, `year`, `publishedAt` (`YYYY-MM-DD`,
  `YYYY-MM` or `YYYY` — as precise as the source), `type` (our vocabulary),
  `typeRaw` (the source's own word).
- **Identifiers:** `doi`, `doiUrl`, `openalexId`, `arxivId`, `arxivVersion`,
  `pmid`, `pmcid`.
- **Access:** `landingUrl`, `pdfUrl`, `isOpenAccess`, `oaStatus`, `license`.
- **Venue:** `venue`, `venueType` (`journal`, `conference`, `repository`,
  `book-series`, `ebook-platform`, `other`), `publisher`, `issn`, `volume`,
  `issue`, `pages`.
- **Impact:** `citationCount`, `referenceCount`, `fwci`, `citationPercentile`,
  `isTop1Percent`, `isTop10Percent`, `isRetracted`, `language`.
- **People:** `authors[]` (`{ name, institutions[], countries[], isGroup }`),
  `authorCount`, `firstAuthor`, `institutions[]`, `institutionCountries[]`.
- **Subject:** `topics[]` (OpenAlex topics, Crossref subjects, arXiv categories
  or PubMed MeSH descriptors), `field`, `subfield`, `domain` (OpenAlex primary
  topic), `keywords[]`, `arxivCategories[]`, `meshTerms[]`,
  `publicationTypes[]`, `journalRef`.
- **Provenance:** `query`, `rank` (1-based position in the source's own result
  list, or in your identifier list), `updatedAt` (when the source last updated
  its record).

Diagnostic `errorType`s, all free: `invalid-id` (not an identifier — no request
made), `invalid-input` (nothing usable to run), `not-found`, `no-results`,
`bad-query` (the source rejected the query, with its reason), `rate-limited`,
`blocked`, `http`, `timeout`, `deadline`, `budget`, `upstream-format`, `skipped`
(Crossref under `openAccessOnly`, or arXiv under a `types` filter without
`preprint`), and `actor-error`.

***

### Which source for what

**OpenAlex** is the default and the one to start with. It indexes more than 250
million works across every discipline and carries, on one record, the abstract,
the citation count with field-normalised impact (`fwci`, percentile, top-1% and
top-10% flags), topics and keywords, the open-access verdict with the best PDF
link, and each author's institutions and countries. It applies every filter this
Actor offers upstream, except the catch-all `other` type.

**Crossref** is where publishers deposit their metadata of record. Use it next to
OpenAlex when you want the publisher's own container title, page range, licence
URL and full-text links, or on its own to resolve DOIs OpenAlex has not picked up
yet. It has no open-access verdict and abstracts only when a publisher deposited
one.

**arXiv** is preprints — physics, mathematics, computer science, statistics,
quantitative biology and finance, economics, electrical engineering. Every paper
has its full abstract and a PDF, and `arxivCategories` for precise topical
searches. There are no citation counts.

**PubMed** is biomedicine and life sciences. Its records carry MeSH descriptors,
author keywords, publication types (so reviews and trials can be told apart) and
structured abstracts; a paper in PubMed Central is marked open access.

### How duplicates are merged

With `dedupe` on (the default) and more than one source in the run, the same work
found by several sources becomes one row. Records are matched on the lower-cased
DOI, then the arXiv id (OpenAlex's arXiv DOIs, `10.48550/arXiv.…`, carry one),
then the PMID. The **base record** is chosen by source — OpenAlex, then Crossref,
then PubMed, then arXiv — never by which answered first, and the base keeps every
value it has; the other records only fill the base's empty columns. `source` names
the base, `sources` names all of them, and the row is charged once. A one-source
run streams its rows as they arrive and drops a paper a second query finds again.

### Authors and privacy

Authors appear **as printed on the paper**: the display name, the names of the
institutions the source matched, and those institutions' country codes. Nothing
else. This Actor never outputs ORCID iDs, author ids, e-mail addresses,
corresponding-author flags, or the raw author-name and raw affiliation strings the
sources also carry (on real records those raw strings contain e-mail addresses).
Free text that we do pass through — an abstract, a publisher-deposited affiliation
line — has any e-mail address replaced with `[e-mail removed]`. A collective author
such as a study group or consortium is marked `isGroup: true`.

Set `includeAuthors: false` to drop the author list and `firstAuthor` entirely;
OpenAlex is then not even asked for them.

### What you are never charged for

- Every diagnostic row — `invalid-id`, `not-found`, `no-results`, `bad-query`,
  `rate-limited`, `blocked`, `http`, `timeout`, `deadline`, `budget`,
  `upstream-format`, `skipped`.
- Duplicates. One paper found by three sources is one row and one charge; a
  paper a second query finds again is dropped, not billed.
- Papers your filters excluded, including Crossref rows dropped by
  `openAccessOnly`. Every filter runs before the row is written.
- The extra requests behind a row: PubMed's `efetch`, the Crossref fallback for a
  DOI OpenAlex does not have, retries.
- A run that returns nothing. A run whose input was usable but produced no paper
  — no match, an unknown identifier, filters that excluded everything, a source
  that refused us or was down, the time budget running out — finishes
  **SUCCEEDED with zero results**, with a status message that says so and points
  at the diagnostic rows, and bills nothing, start fee included. A run finishes
  **FAILED** only when there was nothing usable to attempt (no query and no valid
  identifier) or the Actor itself hit an error.

### Pricing

| Event | FREE | BRONZE | SILVER | GOLD |
|---|---|---|---|---|
| Run started | $0.001 | $0.001 | $0.001 | $0.001 |
| **Paper record** | **$0.0004** | $0.0004 | $0.00032 | $0.00024 |

**$0.40 per 1,000 papers** at the free tier, whichever source a paper came from.
500 papers cost 500 × $0.0004 + $0.001 = **$0.201**.

The run honours your maximum total charge (`ACTOR_MAX_TOTAL_CHARGE_USD`): when it
is reached the Actor stops fetching, leaves a free `budget` row, keeps everything
already delivered and finishes successfully.

***

### Use it from an AI agent, or from code

One JSON object in, one flat array out. The Actor runs with **limited
permissions**, uses **pay-per-event** pricing and never enters Standby, so it
works over the Apify MCP server and with x402 agentic payments.

```bash
curl -X POST "https://api.apify.com/v2/acts/insight.solutions~academic-papers-api/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"queries":["graphene battery"],"maxResults":25,"openAccessOnly":true}'
```

```python
## pip install apify-client
from apify_client import ApifyClient

client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("insight.solutions/academic-papers-api").call(run_input={
    "queries": ["protein structure prediction"],
    "sources": ["openalex", "crossref", "arxiv", "pubmed"],
    "yearFrom": 2020,
    "sortBy": "citations",
    "maxResults": 100,
})

for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    if row["rowType"] != "paper":
        print("skipped:", row["errorType"], row["error"])
        continue
    print(row["year"], row["citationCount"], row["title"], row["doi"], row["pdfUrl"], row["sources"], sep=" | ")
```

***

### FAQ

**Why do Crossref rows usually have no abstract?**
Crossref only has an abstract when the publisher deposited one, and most do not:
1 of the 20 items in our captured Crossref search had one. OpenAlex rebuilds
abstracts for far more works, which is one reason it is the default — and with
`dedupe` on, a paper both sources have comes out with OpenAlex's abstract.

**Why is arXiv slow?**
arXiv's API rules ask for one request every three seconds, one at a time, and this
Actor keeps to that whatever `maxConcurrency` says. Each request returns up to 100
papers, so a 300-paper arXiv walk spends about six seconds waiting.

**Why do citation counts differ between sources?**
Each source counts over its own index. OpenAlex counted 47,662 citations of the
AlphaFold paper where Crossref counted 45,503 on the same day. A merged row keeps
the base record's count (OpenAlex's when OpenAlex has the paper); run with
`dedupe: false` to see both. arXiv and PubMed publish no counts at all.

**How does dedupe choose the base record?**
By source, not by speed: OpenAlex, then Crossref, then PubMed, then arXiv. The
base keeps all of its values and the others fill only its gaps, so the same input
gives the same row every time.

**Can I use my own OpenAlex or NCBI key, or the "polite pool"?**
No. This Actor sends no API key and no e-mail address; it runs in each source's
anonymous pool and paces itself below the published limits.

**Does `minCitations` drop arXiv and PubMed papers?**
No — they have no count to compare, so they are kept with `citationCount: null`.
Use OpenAlex alone when a citation floor must be strict.

### Limitations

- The anonymous pools have limits: OpenAlex allows 100,000 requests a day, NCBI
  three a second, arXiv one every three seconds. A 429 is retried once, after a
  pause and with a fresh proxy session; a second one becomes a free
  `rate-limited` row.
- PubMed's `esearch` does not page past its 10,000th result; narrow the query or
  the years to reach further.
- Crossref has no open-access status, so `openAccessOnly` drops Crossref rows,
  and it does not label review articles, so `types: ["review"]` matches no
  Crossref row.
- A type filter of `other` is checked here rather than upstream, so such a walk
  may read many pages to find a few papers; it stops after ten pages in a row with
  nothing kept.
- No full text: abstracts, metadata and links to the PDF only.
- Merging across sources holds a run's papers until every source has answered,
  up to 50,000 distinct papers per run.
- The upstream formats may change. The parsers are pinned against real responses
  captured on 2026-09-28, and a response that no longer matches becomes a free
  `upstream-format` row rather than a wrong one.

### Sources, terms and attribution

- **OpenAlex** — metadata from OpenAlex (openalex.org), released under the CC0
  public-domain dedication.
- **Crossref** — metadata from the Crossref REST API (api.crossref.org); Crossref
  makes its bibliographic metadata available without restriction, and abstracts
  remain the publishers' own.
- **arXiv** — thank you to arXiv for use of its open access interoperability;
  data is retrieved through the arXiv API under the arXiv API Terms of Use.
- **PubMed** — data from the National Library of Medicine's PubMed via the NCBI
  E-utilities, used under the NCBI usage policy; this Actor is not endorsed by NLM.

### Our other Actors

Every Insight Solutions Actor is pay-per-result with no browser, no login and no API key, and every one of them returns free diagnostic rows instead of billing for failures. Prices are per 1,000 results.

**Video, audio & social**

- [YouTube Transcript API](https://apify.com/insight.solutions/youtube-transcript-api) — captions as timed segments, text, SRT or VTT, with language fallback and translation.
- [YouTube Comments API](https://apify.com/insight.solutions/youtube-comments-api) — comments and replies with likes, pinned and hearted flags, newest or top sort.
- [YouTube Channel API](https://apify.com/insight.solutions/youtube-channel-api) — a channel's videos, Shorts and live streams, plus YouTube search.
- [Podcast Search, Episodes & Charts API](https://apify.com/insight.solutions/podcast-api) — Apple Podcasts search, charts and full episode feeds.
- [Bluesky Scraper](https://apify.com/insight.solutions/bluesky-scraper) — profiles, posts, followers and follows from the public AT Protocol API.
- [Telegram Channel Scraper](https://apify.com/insight.solutions/telegram-channel-scraper) — posts, views and channel stats from public Telegram channels.
- [Substack Scraper](https://apify.com/insight.solutions/substack-scraper) — posts with full free text, comments and publication profiles.
- [Hacker News API](https://apify.com/insight.solutions/hacker-news-api) — stories, comments, users, front page and a structured "Who is hiring?" parser from the official HN APIs.
- [Discourse Forum API](https://apify.com/insight.solutions/discourse-forum-api) — topics, posts and categories from any Discourse community via its own JSON endpoints, usernames only.

**News, documents & the web**

- [Google News Search, Topics & Real Article URLs](https://apify.com/insight.solutions/google-news-api) — news search and topic feeds with the publisher's real URL decoded.
- [Website to Markdown — Content Extractor for LLMs & RAG](https://apify.com/insight.solutions/website-content-extractor) — any site as clean Markdown, text and heading-aware chunks.
- [Internet Archive API](https://apify.com/insight.solutions/internet-archive-api) — archive.org search, item metadata, files and reviews.
- [Wayback Machine Toolkit](https://apify.com/insight.solutions/wayback-toolkit) — archived URL inventories, snapshots and text diffs between dates.
- [Website Technology Detector](https://apify.com/insight.solutions/website-tech-detector) — the tech stack behind any site, with the evidence for each detection.
- [Domain Intelligence API](https://apify.com/insight.solutions/domain-intelligence-api) — DNS, RDAP registration, TLS certificate and HTTP facts in one row per domain.
- [SEO Page Audit](https://apify.com/insight.solutions/seo-page-audit) — sitemap crawl with on-page checks, structured data and broken-link reports.
- [Keyword Suggestions API](https://apify.com/insight.solutions/keyword-suggestions-api) — Google, YouTube, Bing, Amazon and eBay autocomplete with alphabet and question expansions.
- [Website Contact Extractor](https://apify.com/insight.solutions/website-contact-extractor) — emails, phone numbers and social profiles from any list of websites.
- [Web Search Results API](https://apify.com/insight.solutions/web-search-api) — Bing and DuckDuckGo organic results with snippets, no key, no browser.
- [Company Enrichment API](https://apify.com/insight.solutions/company-enrichment-api) — a domain in, a company profile out: firmographics, contacts, tech stack, DNS and hiring signal.
- [Company Dossier API](https://apify.com/insight.solutions/company-dossier-api) — one company in, twelve sections out: profile, tech, contacts, DNS, open roles, news, SEC filings, federal awards, recalls, YC batch and apps.
- [Press Releases API](https://apify.com/insight.solutions/press-releases-api) — GlobeNewswire and PR Newswire releases plus any newsroom feed, by keyword, company, ticker or subject.
- [Federal Register API](https://apify.com/insight.solutions/federal-register-api) — rules, proposed rules, notices and the Public Inspection desk with dockets, comment deadlines and CFR references.
- [RSS & Atom Feed Monitor](https://apify.com/insight.solutions/rss-feed-monitor) — any RSS, Atom or JSON feed (or an OPML file) in, only the new items out, with keyword filters and a webhook.
- [Website Change Monitor](https://apify.com/insight.solutions/website-change-monitor) — watch any pages, diff the text between runs, get change rows with added/removed lines, keyword alerts and a webhook.
- [Wikipedia & Wikidata API](https://apify.com/insight.solutions/wikipedia-api) — article text, search, daily pageviews and Wikidata entity facts, any language edition.

**Business, finance & jobs**

- [Congress & Insider Trades API](https://apify.com/insight.solutions/congress-insider-trades-api) — STOCK Act periodic transaction reports and SEC Form 4 insider trades in one schema.
- [Federal Contracts, Grants & Lobbying API](https://apify.com/insight.solutions/federal-contracts-grants-api) — SAM.gov opportunities, USAspending awards, Grants.gov notices and Senate lobbying filings in one schema.
- [SEC EDGAR API](https://apify.com/insight.solutions/sec-edgar-api) — filings, XBRL financials and full-text search by ticker or CIK.
- [Clinical Trials & FDA API](https://apify.com/insight.solutions/clinical-trials-fda-api) — ClinicalTrials.gov studies plus openFDA recalls, labels, approvals, 510(k)s and adverse-event reports.
- [Product & Vehicle Recalls API](https://apify.com/insight.solutions/product-recalls-api) — CPSC, NHTSA, FDA and USDA recalls, vehicle complaints and ratings, plus a VIN decoder.
- [Y Combinator Companies, Batches & Founders](https://apify.com/insight.solutions/yc-companies-directory) — the YC directory with founders and social links, filterable by batch, industry and hiring status.
- [Career Site Jobs API](https://apify.com/insight.solutions/ats-jobs-api) — jobs straight from Greenhouse, Lever, Ashby, Workable and 10+ other ATS career sites.
- [New Job Postings Monitor](https://apify.com/insight.solutions/job-postings-monitor) — new, closed and changed postings on the career sites you watch.
- [Hiring Signals API — Open Roles & Hiring Surge by Company](https://apify.com/insight.solutions/hiring-signals-api) — one row per company per run: open roles, what opened and closed, department and seniority breakdowns, and a hiring-surge flag.
- [Remote Jobs API](https://apify.com/insight.solutions/remote-jobs-api) — RemoteOK, Remotive, We Work Remotely, Himalayas, Jobicy and more in one schema, deduplicated.
- [Shopify Products API](https://apify.com/insight.solutions/shopify-products-api) — any Shopify store's catalogue, variants, prices and stock signals.
- [Shopify Store Monitor](https://apify.com/insight.solutions/shopify-store-monitor) — price drops, sales, restocks, sell-outs and new products on any Shopify store, one row per change.
- [Public Tenders API](https://apify.com/insight.solutions/public-tenders-api) — EU TED, UK Find a Tender and Contracts Finder notices by keyword, CPV code, country, stage and deadline.
- [Nonprofit & IRS 990 Lookup API](https://apify.com/insight.solutions/nonprofit-990-api) — search US nonprofits and get EIN, NTEE code and multi-year Form 990 financials.
- [OpenStreetMap Places API](https://apify.com/insight.solutions/osm-places-api) — businesses and points of interest by category and area from OpenStreetMap: name, address, coordinates, website, phone, opening hours.

**Apps & games**

- [App Store & Google Play Reviews API](https://apify.com/insight.solutions/app-reviews-api) — reviews from both stores with ratings, versions and developer replies.
- [App Store Top Charts & App Search API](https://apify.com/insight.solutions/app-charts-api) — Apple top charts by country and genre, plus app search and details.
- [App Store Keyword Rank Tracker](https://apify.com/insight.solutions/app-store-keyword-rank-tracker) — where any app ranks for any keyword on the App Store and Google Play, with rank changes and ASO suggestions.
- [Steam Reviews API](https://apify.com/insight.solutions/steam-reviews-api) — Steam reviews with playtime, helpfulness and game details.
- [Steam Game Data API](https://apify.com/insight.solutions/steam-store-stats-api) — prices, tags, review scores, live player counts and top charts.

# Actor input Schema

## `queries` (type: `array`):

Free-text searches. Each query is searched once in every source you pick below, and each of those walks returns up to `maxResults` papers. Leave empty to only look up the identifiers below.

## `sources` (type: `array`):

Where the queries are searched. OpenAlex is the broadest (abstracts, citations, topics, open-access links, institutions). Crossref is the publishers' own metadata of record. arXiv is preprints with full abstracts and PDFs. PubMed is biomedicine, with MeSH terms and abstracts. Identifier lookups always go to the source that owns the identifier, whatever this says.

## `maxResults` (type: `integer`):

How many papers each query may return from each source. `0` means no cap (the time budget still applies). Identifier lookups are not capped: they return what you asked for.

## `dois` (type: `array`):

Direct lookups, in any form: `10.1038/s41586-021-03819-2`, `doi:10.1038/…` or `https://doi.org/10.1038/…`. Resolved through OpenAlex first, then Crossref when OpenAlex has no record. An unknown DOI is a free `not-found` row.

## `arxivIds` (type: `array`):

`2303.08774`, `2303.08774v2`, `arXiv:2303.08774` or `https://arxiv.org/abs/2303.08774`; the pre-2007 form `hep-th/9901001` works too. Looked up on arXiv, 50 per request.

## `pmids` (type: `array`):

PMIDs such as `38253521`, `PMID: 38253521` or a pubmed.ncbi.nlm.nih.gov link. Looked up on PubMed, 50 per request, with abstracts, MeSH terms and keywords.

## `arxivCategories` (type: `array`):

Limit the arXiv search to categories such as `cs.CL`, `cs.LG` or `q-bio.NC` (several are ORed). Setting one adds arXiv to the sources. With no search query at all, the categories are browsed on their own, newest first when `sortBy` is `date`.

## `yearFrom` (type: `integer`):

Earliest publication year, inclusive. Applied upstream by OpenAlex, Crossref and PubMed, and to arXiv's first-submission date here.

## `yearTo` (type: `integer`):

Latest publication year, inclusive.

## `types` (type: `array`):

Keep only these kinds of work. Leave empty for all. arXiv holds preprints only; Crossref does not label review articles, so `review` never matches a Crossref record; PubMed types come from its publication types.

## `openAccessOnly` (type: `boolean`):

Keep only papers that are free to read. OpenAlex applies this upstream; every arXiv paper qualifies; PubMed papers qualify when they are in PubMed Central. Crossref publishes no open-access status, so Crossref results are dropped (one free `skipped` row says how many).

## `minCitations` (type: `integer`):

Keep only papers cited at least this many times. Applied upstream by OpenAlex and to Crossref's own count here. arXiv and PubMed publish no citation counts, so their papers are kept with `citationCount: null`.

## `sortBy` (type: `string`):

`relevance` (each source's own ranking), `citations` (most cited first — OpenAlex and Crossref; arXiv and PubMed fall back to relevance) or `date` (newest first, in every source).

## `includeAbstract` (type: `boolean`):

Fill the `abstract` column. OpenAlex's is rebuilt from its word index; Crossref has one only when the publisher deposited it; arXiv's is always there; PubMed's needs one extra request per 50 papers. Turn off for smaller, faster rows.

## `includeAuthors` (type: `boolean`):

Fill `authors` with each author's name as printed on the paper, their institutions and those institutions' countries — nothing else, never an ORCID or an e-mail address. Turn off to drop the author list and `firstAuthor` entirely.

## `dedupe` (type: `boolean`):

One paper, one row: records from several sources are merged on DOI, then arXiv id, then PMID. The OpenAlex record is the base and the others fill its empty columns; `sources` lists every source that had it. You pay for the merged row once. Turn off to get (and pay for) every source's row separately.

## `maxConcurrency` (type: `integer`):

How many searches and lookups run at once. arXiv is always one request at a time, three seconds apart, and PubMed at most three a second, whatever this is set to — those are their published rules.

## `maxRunSecs` (type: `integer`):

Stop fetching after this many seconds and keep everything already returned.

## `proxyConfiguration` (type: `object`):

All four sources answered from Apify's datacenter proxy and from no proxy at all in testing, so the cheap datacenter pool is the default and a cleared field still works.

## Actor input object example

```json
{
  "queries": [
    "graphene battery"
  ],
  "sources": [
    "openalex"
  ],
  "maxResults": 10,
  "dois": [],
  "arxivIds": [],
  "pmids": [],
  "arxivCategories": [],
  "types": [],
  "openAccessOnly": false,
  "sortBy": "relevance",
  "includeAbstract": true,
  "includeAuthors": true,
  "dedupe": true,
  "maxConcurrency": 3,
  "maxRunSecs": 240,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

One row per paper, merged across the sources that had it, plus free diagnostic rows. Delivered as JSON items in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "graphene battery"
    ],
    "sources": [
        "openalex"
    ],
    "maxResults": 10,
    "sortBy": "relevance",
    "includeAbstract": true,
    "includeAuthors": true,
    "dedupe": true,
    "maxConcurrency": 3,
    "maxRunSecs": 240,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("insight.solutions/academic-papers-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["graphene battery"],
    "sources": ["openalex"],
    "maxResults": 10,
    "sortBy": "relevance",
    "includeAbstract": True,
    "includeAuthors": True,
    "dedupe": True,
    "maxConcurrency": 3,
    "maxRunSecs": 240,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("insight.solutions/academic-papers-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "graphene battery"
  ],
  "sources": [
    "openalex"
  ],
  "maxResults": 10,
  "sortBy": "relevance",
  "includeAbstract": true,
  "includeAuthors": true,
  "dedupe": true,
  "maxConcurrency": 3,
  "maxRunSecs": 240,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call insight.solutions/academic-papers-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,insight.solutions/academic-papers-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/gudft4PKjkVup0wtL/builds/1vnLZ4GtOBvTqmaeJ/openapi.json
