# arXiv Preprints API (`conserving_celerytop/preprint-papers-api`) Actor

Search arXiv, bioRxiv and medRxiv preprints, or look up arXiv IDs and DOIs. One row per preprint with title, abstract, categories, version, dates, links and licence. No author fields. Independent tool, not affiliated with arXiv, bioRxiv or medRxiv.

- **URL**: https://apify.com/conserving_celerytop/preprint-papers-api.md
- **Developed by:** [Don Mangu](https://apify.com/conserving_celerytop) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.40 / 1,000 preprint records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

Use this preprint API to search arXiv, bioRxiv and medRxiv, or look up arXiv IDs and bioRxiv and medRxiv DOIs. Get the title, abstract, categories, version, dates, links and licence for $2.00 per 1,000 preprints.

Type a search such as "retrieval augmented generation", pick a category or a date range, or paste a list of arXiv IDs and DOIs. Each result is one preprint. The Actor does not return author, affiliation or e-mail fields. It also removes e-mail addresses that authors write inside a title or abstract. You need no login, no API key and no proxy.

This is an unofficial tool built by an independent developer. It is not affiliated with, endorsed by or operated by arXiv, Cornell University, bioRxiv, medRxiv or Cold Spring Harbor Laboratory. Thank you to arXiv for use of its open access interoperability.

### Sample output

One row per preprint. This is a real row from a run on 7 October 2026:

```json
{
  "id": "hep-th/9901001",
  "server": "arxiv",
  "doi": "10.48550/arXiv.hep-th/9901001",
  "publishedDoi": "10.1143/PTP.101.1155",
  "title": "String Junctions and Their Duals in Heterotic String Theory",
  "abstract": "We explicitly give the correspondence between spectra of heterotic string theory compactified on $T^2$ and string junctions in type IIB theory compactified on $S^2$.",
  "categories": [
    "hep-th"
  ],
  "primaryCategory": "hep-th",
  "version": 3,
  "firstPostedDate": "1999-01-01",
  "updatedDate": "1999-05-10",
  "publicationType": null,
  "licence": null,
  "absUrl": "https://arxiv.org/abs/hep-th/9901001",
  "pdfUrl": "https://arxiv.org/pdf/hep-th/9901001",
  "query": "hep-th/9901001",
  "matchedBy": "arxiv_id",
  "retrievedAt": "2026-10-07T11:58:47.708Z",
  "dataSource": "arXiv (arxiv.org)",
  "resultStatus": "ok"
}
```

### How to get preprint data with this Actor

1. Click **Try for free**. No API key is needed.
2. In **Search queries**, enter one search per line. Use plain words, quotes for a phrase, and AND, OR or ANDNOT. On arXiv you can also use field prefixes such as `ti:`, `abs:` and `cat:`.
3. Or enter arXiv IDs, arXiv links, or bioRxiv and medRxiv DOIs in **arXiv ids or DOIs**, one per line. You can mix them. A link to the preprint page also works.
4. Under **Servers**, choose arXiv, bioRxiv, medRxiv, or any mix.
5. Optionally set **Date from**, **Date to**, **arXiv categories**, **bioRxiv and medRxiv subject areas** and **Sort by**. A preprint must pass every filter you fill in.
6. Set **Max results**. The default is 100 and the limit is 10,000 rows per run.
7. Click **Start**, open the **Overview** view, and export as JSON, CSV or Excel.

Typical uses:

- Build a reading list or a literature table for a topic, newest first.
- Turn a column of arXiv IDs or DOIs in a sheet into titles, dates and abstracts.
- Watch one arXiv category every day and send the new preprints to a team.
- Check whether a preprint has a journal version, and see the DOI of that version.
- Feed abstracts to a classifier, a search index or a language model.

### What you get

| Field | Meaning |
|---|---|
| id, server | arXiv ID without version (for example 2101.00001), or the bioRxiv or medRxiv DOI. Server is arxiv, biorxiv or medrxiv |
| doi, publishedDoi | DOI of the preprint (for arXiv, 10.48550/arXiv.ID), and the DOI of the journal version when the source lists one |
| title, abstract | Plain text. abstract is null when the preprint has none or Include abstract is off. LaTeX in arXiv text stays as written |
| categories, primaryCategory | arXiv category codes such as cs.LG, or the bioRxiv or medRxiv subject area |
| version | Version number of the listed record. arXiv rows show the latest version |
| firstPostedDate, updatedDate | Date of the first version, and date of the listed version (YYYY-MM-DD). For bioRxiv and medRxiv the first date is set only when the listed record is version 1 |
| publicationType | bioRxiv and medRxiv article type, for example new results. Null for arXiv |
| licence | bioRxiv and medRxiv licence of the preprint, as the API writes it. Null for arXiv |
| absUrl, pdfUrl | Links to the preprint page and to the PDF on the source site. The Actor does not download PDFs |
| query, matchedBy | The search or identifier that produced the row, and search, arxiv_id or doi |
| retrievedAt, dataSource, resultStatus | Read time, the source, and ok or not_found |

The Actor does not return author fields, affiliations, e-mail fields, ORCID iDs or the free-text comment and journal reference fields of arXiv. It does not return PDFs or full text.

### Pricing

Pay per event: **$2.00 per 1,000 preprint rows** on the Free plan, with lower prices on Bronze ($1.80), Silver ($1.60) and Gold ($1.40) plans. The start of a run costs $0.00005. An identifier that matches no preprint returns one not_found row and counts as one result, because the lookup was done. A search with no hits returns nothing and costs nothing.

Worked example: looking up 5 arXiv IDs returns 5 rows and costs 5 x $0.002 = $0.01, plus the start fee. A search that returns 1,000 preprints costs 1,000 x $0.002 = $2.00.

arXiv asks for no more than one request every three seconds, so a run is paced. Two runs of 1,000 rows took 28.6 and 28.9 seconds on my computer and made 10 requests each. A run of 3,000 rows took 90.9 seconds and made 30 requests. Rows saved before a stop are charged, and rows that were not saved are not.

### Input examples

Search a topic on arXiv, newest first, from 2025:

```json
{
  "queries": ["retrieval augmented generation"],
  "servers": ["arxiv"],
  "dateFrom": "2025-01-01",
  "sort": "newest",
  "maxResults": 100
}
```

On 7 October 2026 arXiv reported 5,989 matches for that search without the sort, and the run returned the 100 newest. In those 100 rows, 75 contain the exact words retrieval, augmented and generation in the title or abstract, and 91 contain the stems retriev, augment and generat. arXiv matches word stems, so a search for "generation" also finds related word forms.

List one category for one month, without keywords:

```json
{
  "queries": [],
  "servers": ["arxiv"],
  "arxivCategories": ["q-bio.NC"],
  "dateFrom": "2025-09-01",
  "dateTo": "2025-09-30",
  "includeAbstract": false
}
```

arXiv reported 124 matches for q-bio.NC in September 2025. The run returned all 124 in 2 requests and then stopped.

Look up papers by identifier:

```json
{
  "ids": ["1706.03762", "hep-th/9901001", "10.1101/2020.03.01.972935"]
}
```

A real lookup of five well known arXiv IDs on 7 October 2026 returned:

| arXiv ID | Version | First posted | Category | Title |
|---|---|---|---|---|
| 1706.03762 | 7 | 2017-06-12 | cs.CL | Attention Is All You Need |
| 1810.04805 | 2 | 2018-10-11 | cs.CL | BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding |
| 2005.14165 | 4 | 2020-05-28 | cs.CL | Language Models are Few-Shot Learners |
| 1512.03385 | 1 | 2015-12-10 | cs.CV | Deep Residual Learning for Image Recognition |
| hep-th/9901001 | 3 | 1999-01-01 | hep-th | String Junctions and Their Duals in Heterotic String Theory |

Latest bioRxiv and medRxiv preprints on a topic, last 30 days:

```json
{
  "queries": ["gut microbiome"],
  "servers": ["biorxiv", "medrxiv"]
}
```

### Output example for an identifier that matches nothing

A real row for the arXiv ID 9999.99999:

```json
{
  "id": null,
  "server": null,
  "doi": null,
  "publishedDoi": null,
  "title": null,
  "abstract": null,
  "categories": null,
  "primaryCategory": null,
  "version": null,
  "firstPostedDate": null,
  "updatedDate": null,
  "publicationType": null,
  "licence": null,
  "absUrl": null,
  "pdfUrl": null,
  "query": "9999.99999",
  "matchedBy": null,
  "retrievedAt": "2026-10-07T11:59:22.842Z",
  "dataSource": "arXiv (arxiv.org)",
  "resultStatus": "not_found"
}
```

### Run it every day

New preprints appear every day. In the Apify Console, open **Schedules**, create a schedule with the cron expression `0 7 * * *`, and add this Actor with your category or search and a **Date from** of the previous day. Use a dataset export, a webhook, or an integration with Make, Zapier or n8n to send the rows on. You can also call the Actor from your own code with the Apify API.

### FAQ

**Where does the data come from?** From the public arXiv API (export.arxiv.org) and the public bioRxiv and medRxiv API (api.biorxiv.org). The Actor does not scrape web pages and does not use a key.

**Is it legal to use?** The rules differ by source, and this is not legal advice. arXiv states in its API terms of use: "You are free to use descriptive metadata about arXiv e-prints under the terms of the Creative Commons Universal (CC0 1.0) Public Domain Declaration." The same terms say that you may not store and serve the e-prints themselves (PDFs and source files) without permission of the copyright holder, so the Actor returns only links to them. arXiv asks users to acknowledge it: Thank you to arXiv for use of its open access interoperability. For bioRxiv and medRxiv, the authors keep the copyright of their text and choose a licence, which can be CC BY, CC BY-NC, CC0 or a stricter one. The licence of each preprint is in the `licence` field. Check it before you republish an abstract, and switch the abstract off with **Include abstract** if you need metadata only.

**Why are there no authors?** Author names, affiliations and e-mail addresses are personal data. The Actor does not read the author fields of either source, so the rows are lower risk under privacy rules. It also does not read the arXiv comment and journal reference fields. In a test, one real journal reference started with an author name, so those fields are left out. The title and abstract are free text and can still contain a name. The Actor removes e-mail addresses that are written inside title and abstract text.

**How does a search work on arXiv?** Your words must appear in the title or the abstract. Each word is matched with the arXiv search engine, which also matches word stems. Choose Everywhere in **Search in** to search all arXiv fields. A query that names a field, such as `ti:"diffusion models" AND cat:cs.CV`, is used as you typed it. A search by author name works, but the author names are not in the output.

**How does a search work on bioRxiv and medRxiv?** Their API lists preprints by date and has no keyword search. The Actor reads your date window one week at a time, newest week first, and keeps the rows whose title or abstract match your words as whole words. AND, OR, ANDNOT and quotes work. Field prefixes such as `cat:` work on arXiv only, and a search that uses them is skipped for bioRxiv and medRxiv. If you give no dates, the Actor reads the last 30 days. The longest window is 366 days, because a long window means many requests. The Actor makes at most two requests per second to this API. This part is new. If a run differs from the API documentation, please tell me in the Issues tab.

**What does a bioRxiv or medRxiv row show when a preprint has several versions?** The latest version in your window. The `version` field shows which one. A lookup by DOI returns the latest version overall.

**How do identifiers match?** An arXiv ID can be new style (2101.00001) or old style (hep-th/9901001), with or without a version, as a link, or as the DOI 10.48550/arXiv.ID. A version number is ignored and the latest version is returned. A DOI that starts with 10.1101/ is looked up on bioRxiv and then on medRxiv. A DOI of a journal is not accepted. If two identifiers point to the same preprint, it is returned once.

**What are the limits?** Up to 50 searches, 1,000 identifiers and 10,000 rows per run. Max results counts all rows, including not_found rows. **Max results per search** limits one search on one server and is off by default.

**How complete is the data?** In a run of 1,000 rows on 7 October 2026 (category cs.LG, September and October 2025), 100% of rows had an abstract, 81.6% listed more than one category and 7.2% had a journal DOI. Journal DOIs appear only when the authors gave one to arXiv.

**What if the run fails?** If a source is down or sends a reply the Actor cannot read, the run stops with a clear message. Rows saved before the stop stay in the dataset. If you chose several servers and one fails, the others still run, and the run ends with an error so that you notice. Run it again later.

### Related Actors

Other Actors by the same author, found on the author's Store profile:

- Crossref Works Search: papers, DOIs and citations from Crossref (https://apify.com/conserving_celerytop/crossref-works-search).
- PubMed Papers API: PubMed and Europe PMC records without author fields (Store address to be added after it is published).

I built this Actor myself as an independent developer. Data: arXiv (arxiv.org), bioRxiv (biorxiv.org) and medRxiv (medrxiv.org). Thank you to arXiv for use of its open access interoperability.

# Actor input Schema

## `queries` (type: `array`):

Enter one search per line, in plain words, for example: retrieval augmented generation, or "graph neural network" AND molecules. AND, OR and ANDNOT work, and quotes mark a phrase. On arXiv you can also use field prefixes such as ti:, abs: and cat:. Leave empty to list preprints by category or date.

## `servers` (type: `array`):

Which servers to search. Identifier lookups try the right server by themselves.

## `searchIn` (type: `string`):

Where plain search words must appear on arXiv. Title and abstract gives on-topic results. bioRxiv and medRxiv are always matched in title and abstract.

## `ids` (type: `array`):

Enter one identifier per line: an arXiv id such as 1706.03762 or hep-th/9901001, an arXiv link, or a bioRxiv or medRxiv DOI such as 10.1101/2020.03.01.972935. A link to the preprint page also works. An identifier that matches nothing returns one not_found row. Filters do not apply to lookups.

## `dateFrom` (type: `string`):

First day, as YYYY-MM-DD, for example 2025-01-31. On arXiv it filters the date of the first version. bioRxiv and medRxiv read the last 30 days when you give no dates.

## `dateTo` (type: `string`):

Last day, as YYYY-MM-DD. bioRxiv and medRxiv accept a window of at most 366 days.

## `arxivCategories` (type: `array`):

Keep arXiv preprints in these categories, one per line, for example cs.LG, hep-th or math.\*. Applies to arXiv only.

## `biorxivCategories` (type: `array`):

Keep bioRxiv and medRxiv preprints in these subject areas, one per line, for example Neuroscience or Infectious Diseases. Applies to bioRxiv and medRxiv only.

## `sort` (type: `string`):

Order of arXiv results. bioRxiv and medRxiv rows always come newest first, because their API has no ranking.

## `includeAbstract` (type: `boolean`):

Add the abstract text to each row. arXiv states that its descriptive metadata is free to use under CC0. bioRxiv and medRxiv authors keep the copyright of their text, and each row shows the licence. Switch this off if you only need metadata.

## `abstractMaxChars` (type: `integer`):

Cut each abstract to this many characters. Use 0 for the full abstract.

## `maxPerSearch` (type: `integer`):

Stop one search on one server after this many rows. Each query and each server counts as one search. Leave empty for no extra cap: only Max results applies.

## `maxResults` (type: `integer`):

Return at most this many rows in total, across all searches and lookups, including not_found rows. Each row returned is one charged result.

## Actor input object example

```json
{
  "queries": [
    "retrieval augmented generation"
  ],
  "servers": [
    "arxiv"
  ],
  "searchIn": "title_abstract",
  "sort": "relevance",
  "includeAbstract": true,
  "abstractMaxChars": 0,
  "maxResults": 25
}
```

# Actor output Schema

## `preprints` (type: `string`):

No description

## `stats` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "retrieval augmented generation"
    ],
    "servers": [
        "arxiv"
    ],
    "maxResults": 25
};

// Run the Actor and wait for it to finish
const run = await client.actor("conserving_celerytop/preprint-papers-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["retrieval augmented generation"],
    "servers": ["arxiv"],
    "maxResults": 25,
}

# Run the Actor and wait for it to finish
run = client.actor("conserving_celerytop/preprint-papers-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "retrieval augmented generation"
  ],
  "servers": [
    "arxiv"
  ],
  "maxResults": 25
}' |
apify call conserving_celerytop/preprint-papers-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,conserving_celerytop/preprint-papers-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/K2RyWQ65npQTHd7em/builds/UOVxW7yvKDBuVkuRg/openapi.json
