arXiv Preprints API
Pricing
from $1.40 / 1,000 preprint records
arXiv Preprints API
Search arXiv, bioRxiv and medRxiv preprints, or look up arXiv IDs and DOIs. One row per preprint with title, abstract, categories, version, dates, links and licence. No author fields. Independent tool, not affiliated with arXiv, bioRxiv or medRxiv.
Pricing
from $1.40 / 1,000 preprint records
Rating
0.0
(0)
Developer
Don Mangu
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 hours ago
Last modified
Categories
Share
Use this preprint API to search arXiv, bioRxiv and medRxiv, or look up arXiv IDs and bioRxiv and medRxiv DOIs. Get the title, abstract, categories, version, dates, links and licence for $2.00 per 1,000 preprints.
Type a search such as "retrieval augmented generation", pick a category or a date range, or paste a list of arXiv IDs and DOIs. Each result is one preprint. The Actor does not return author, affiliation or e-mail fields. It also removes e-mail addresses that authors write inside a title or abstract. You need no login, no API key and no proxy.
This is an unofficial tool built by an independent developer. It is not affiliated with, endorsed by or operated by arXiv, Cornell University, bioRxiv, medRxiv or Cold Spring Harbor Laboratory. Thank you to arXiv for use of its open access interoperability.
Sample output
One row per preprint. This is a real row from a run on 7 October 2026:
{"id": "hep-th/9901001","server": "arxiv","doi": "10.48550/arXiv.hep-th/9901001","publishedDoi": "10.1143/PTP.101.1155","title": "String Junctions and Their Duals in Heterotic String Theory","abstract": "We explicitly give the correspondence between spectra of heterotic string theory compactified on $T^2$ and string junctions in type IIB theory compactified on $S^2$.","categories": ["hep-th"],"primaryCategory": "hep-th","version": 3,"firstPostedDate": "1999-01-01","updatedDate": "1999-05-10","publicationType": null,"licence": null,"absUrl": "https://arxiv.org/abs/hep-th/9901001","pdfUrl": "https://arxiv.org/pdf/hep-th/9901001","query": "hep-th/9901001","matchedBy": "arxiv_id","retrievedAt": "2026-10-07T11:58:47.708Z","dataSource": "arXiv (arxiv.org)","resultStatus": "ok"}
How to get preprint data with this Actor
- Click Try for free. No API key is needed.
- In Search queries, enter one search per line. Use plain words, quotes for a phrase, and AND, OR or ANDNOT. On arXiv you can also use field prefixes such as
ti:,abs:andcat:. - Or enter arXiv IDs, arXiv links, or bioRxiv and medRxiv DOIs in arXiv ids or DOIs, one per line. You can mix them. A link to the preprint page also works.
- Under Servers, choose arXiv, bioRxiv, medRxiv, or any mix.
- Optionally set Date from, Date to, arXiv categories, bioRxiv and medRxiv subject areas and Sort by. A preprint must pass every filter you fill in.
- Set Max results. The default is 100 and the limit is 10,000 rows per run.
- Click Start, open the Overview view, and export as JSON, CSV or Excel.
Typical uses:
- Build a reading list or a literature table for a topic, newest first.
- Turn a column of arXiv IDs or DOIs in a sheet into titles, dates and abstracts.
- Watch one arXiv category every day and send the new preprints to a team.
- Check whether a preprint has a journal version, and see the DOI of that version.
- Feed abstracts to a classifier, a search index or a language model.
What you get
| Field | Meaning |
|---|---|
| id, server | arXiv ID without version (for example 2101.00001), or the bioRxiv or medRxiv DOI. Server is arxiv, biorxiv or medrxiv |
| doi, publishedDoi | DOI of the preprint (for arXiv, 10.48550/arXiv.ID), and the DOI of the journal version when the source lists one |
| title, abstract | Plain text. abstract is null when the preprint has none or Include abstract is off. LaTeX in arXiv text stays as written |
| categories, primaryCategory | arXiv category codes such as cs.LG, or the bioRxiv or medRxiv subject area |
| version | Version number of the listed record. arXiv rows show the latest version |
| firstPostedDate, updatedDate | Date of the first version, and date of the listed version (YYYY-MM-DD). For bioRxiv and medRxiv the first date is set only when the listed record is version 1 |
| publicationType | bioRxiv and medRxiv article type, for example new results. Null for arXiv |
| licence | bioRxiv and medRxiv licence of the preprint, as the API writes it. Null for arXiv |
| absUrl, pdfUrl | Links to the preprint page and to the PDF on the source site. The Actor does not download PDFs |
| query, matchedBy | The search or identifier that produced the row, and search, arxiv_id or doi |
| retrievedAt, dataSource, resultStatus | Read time, the source, and ok or not_found |
The Actor does not return author fields, affiliations, e-mail fields, ORCID iDs or the free-text comment and journal reference fields of arXiv. It does not return PDFs or full text.
Pricing
Pay per event: $2.00 per 1,000 preprint rows on the Free plan, with lower prices on Bronze ($1.80), Silver ($1.60) and Gold ($1.40) plans. The start of a run costs $0.00005. An identifier that matches no preprint returns one not_found row and counts as one result, because the lookup was done. A search with no hits returns nothing and costs nothing.
Worked example: looking up 5 arXiv IDs returns 5 rows and costs 5 x $0.002 = $0.01, plus the start fee. A search that returns 1,000 preprints costs 1,000 x $0.002 = $2.00.
arXiv asks for no more than one request every three seconds, so a run is paced. Two runs of 1,000 rows took 28.6 and 28.9 seconds on my computer and made 10 requests each. A run of 3,000 rows took 90.9 seconds and made 30 requests. Rows saved before a stop are charged, and rows that were not saved are not.
Input examples
Search a topic on arXiv, newest first, from 2025:
{"queries": ["retrieval augmented generation"],"servers": ["arxiv"],"dateFrom": "2025-01-01","sort": "newest","maxResults": 100}
On 7 October 2026 arXiv reported 5,989 matches for that search without the sort, and the run returned the 100 newest. In those 100 rows, 75 contain the exact words retrieval, augmented and generation in the title or abstract, and 91 contain the stems retriev, augment and generat. arXiv matches word stems, so a search for "generation" also finds related word forms.
List one category for one month, without keywords:
{"queries": [],"servers": ["arxiv"],"arxivCategories": ["q-bio.NC"],"dateFrom": "2025-09-01","dateTo": "2025-09-30","includeAbstract": false}
arXiv reported 124 matches for q-bio.NC in September 2025. The run returned all 124 in 2 requests and then stopped.
Look up papers by identifier:
{"ids": ["1706.03762", "hep-th/9901001", "10.1101/2020.03.01.972935"]}
A real lookup of five well known arXiv IDs on 7 October 2026 returned:
| arXiv ID | Version | First posted | Category | Title |
|---|---|---|---|---|
| 1706.03762 | 7 | 2017-06-12 | cs.CL | Attention Is All You Need |
| 1810.04805 | 2 | 2018-10-11 | cs.CL | BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding |
| 2005.14165 | 4 | 2020-05-28 | cs.CL | Language Models are Few-Shot Learners |
| 1512.03385 | 1 | 2015-12-10 | cs.CV | Deep Residual Learning for Image Recognition |
| hep-th/9901001 | 3 | 1999-01-01 | hep-th | String Junctions and Their Duals in Heterotic String Theory |
Latest bioRxiv and medRxiv preprints on a topic, last 30 days:
{"queries": ["gut microbiome"],"servers": ["biorxiv", "medrxiv"]}
Output example for an identifier that matches nothing
A real row for the arXiv ID 9999.99999:
{"id": null,"server": null,"doi": null,"publishedDoi": null,"title": null,"abstract": null,"categories": null,"primaryCategory": null,"version": null,"firstPostedDate": null,"updatedDate": null,"publicationType": null,"licence": null,"absUrl": null,"pdfUrl": null,"query": "9999.99999","matchedBy": null,"retrievedAt": "2026-10-07T11:59:22.842Z","dataSource": "arXiv (arxiv.org)","resultStatus": "not_found"}
Run it every day
New preprints appear every day. In the Apify Console, open Schedules, create a schedule with the cron expression 0 7 * * *, and add this Actor with your category or search and a Date from of the previous day. Use a dataset export, a webhook, or an integration with Make, Zapier or n8n to send the rows on. You can also call the Actor from your own code with the Apify API.
FAQ
Where does the data come from? From the public arXiv API (export.arxiv.org) and the public bioRxiv and medRxiv API (api.biorxiv.org). The Actor does not scrape web pages and does not use a key.
Is it legal to use? The rules differ by source, and this is not legal advice. arXiv states in its API terms of use: "You are free to use descriptive metadata about arXiv e-prints under the terms of the Creative Commons Universal (CC0 1.0) Public Domain Declaration." The same terms say that you may not store and serve the e-prints themselves (PDFs and source files) without permission of the copyright holder, so the Actor returns only links to them. arXiv asks users to acknowledge it: Thank you to arXiv for use of its open access interoperability. For bioRxiv and medRxiv, the authors keep the copyright of their text and choose a licence, which can be CC BY, CC BY-NC, CC0 or a stricter one. The licence of each preprint is in the licence field. Check it before you republish an abstract, and switch the abstract off with Include abstract if you need metadata only.
Why are there no authors? Author names, affiliations and e-mail addresses are personal data. The Actor does not read the author fields of either source, so the rows are lower risk under privacy rules. It also does not read the arXiv comment and journal reference fields. In a test, one real journal reference started with an author name, so those fields are left out. The title and abstract are free text and can still contain a name. The Actor removes e-mail addresses that are written inside title and abstract text.
How does a search work on arXiv? Your words must appear in the title or the abstract. Each word is matched with the arXiv search engine, which also matches word stems. Choose Everywhere in Search in to search all arXiv fields. A query that names a field, such as ti:"diffusion models" AND cat:cs.CV, is used as you typed it. A search by author name works, but the author names are not in the output.
How does a search work on bioRxiv and medRxiv? Their API lists preprints by date and has no keyword search. The Actor reads your date window one week at a time, newest week first, and keeps the rows whose title or abstract match your words as whole words. AND, OR, ANDNOT and quotes work. Field prefixes such as cat: work on arXiv only, and a search that uses them is skipped for bioRxiv and medRxiv. If you give no dates, the Actor reads the last 30 days. The longest window is 366 days, because a long window means many requests. The Actor makes at most two requests per second to this API. This part is new. If a run differs from the API documentation, please tell me in the Issues tab.
What does a bioRxiv or medRxiv row show when a preprint has several versions? The latest version in your window. The version field shows which one. A lookup by DOI returns the latest version overall.
How do identifiers match? An arXiv ID can be new style (2101.00001) or old style (hep-th/9901001), with or without a version, as a link, or as the DOI 10.48550/arXiv.ID. A version number is ignored and the latest version is returned. A DOI that starts with 10.1101/ is looked up on bioRxiv and then on medRxiv. A DOI of a journal is not accepted. If two identifiers point to the same preprint, it is returned once.
What are the limits? Up to 50 searches, 1,000 identifiers and 10,000 rows per run. Max results counts all rows, including not_found rows. Max results per search limits one search on one server and is off by default.
How complete is the data? In a run of 1,000 rows on 7 October 2026 (category cs.LG, September and October 2025), 100% of rows had an abstract, 81.6% listed more than one category and 7.2% had a journal DOI. Journal DOIs appear only when the authors gave one to arXiv.
What if the run fails? If a source is down or sends a reply the Actor cannot read, the run stops with a clear message. Rows saved before the stop stay in the dataset. If you chose several servers and one fails, the others still run, and the run ends with an error so that you notice. Run it again later.
Related Actors
Other Actors by the same author, found on the author's Store profile:
- Crossref Works Search: papers, DOIs and citations from Crossref (https://apify.com/conserving_celerytop/crossref-works-search).
- PubMed Papers API: PubMed and Europe PMC records without author fields (Store address to be added after it is published).
I built this Actor myself as an independent developer. Data: arXiv (arxiv.org), bioRxiv (biorxiv.org) and medRxiv (medrxiv.org). Thank you to arXiv for use of its open access interoperability.