arXiv Scraper - Research Papers & Abstracts
Pricing
from $1.00 / 1,000 results
arXiv Scraper - Research Papers & Abstracts
Extract research papers from arXiv by keyword, category, author or date. Returns arXiv ID, title, full abstract, all authors, subject categories, publication and revision date plus links to the abstract page and the PDF. For AI research tracking, literature reviews and competitive R&D intelligence.
Pricing
from $1.00 / 1,000 results
Rating
0.0
(0)
Developer
Ryan Zinburg
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
0
Monthly active users
21 hours ago
Last modified
Categories
Share
arXiv Scraper - Research Papers, Abstracts & Authors
Extract paper metadata from arXiv, the preprint server where most machine learning, physics and mathematics research appears first, typically months before journal publication. Search by keyword, category, author or date window and export a complete abstract corpus.
Official arXiv API, no key required, no proxy needed.
What you get per paper
| Field | Example |
|---|---|
id | 2609.01234v1 |
title | Scaling Laws for Sparse Mixture-of-Experts Models |
abstract | the full abstract text |
authors | every author, in order |
categories | ["cs.LG", "cs.CL", "stat.ML"] |
publishedDate | first submission date |
updatedDate | latest revision date |
arxivUrl | abstract page |
pdfUrl | direct PDF link |
Search filters
- searchQuery - free text across title and abstract, e.g.
large language models,diffusion - categories - arXiv categories, e.g.
cs.LG,cs.CL,cs.CV,stat.ML,math.OC - dateFrom - earliest submission date
- sortBy - relevance, newest submissions or latest updates
- maxResults - how many papers to save
Example input
{"searchQuery": "large language models","categories": "cs.LG,cs.CL","sortBy": "submittedDate","maxResults": 200}
Use cases
- AI research tracking - follow a subfield week by week without reading every listing by hand
- Literature reviews - export an abstract corpus and search or summarise it offline
- Competitive R&D intelligence - see which labs and authors publish in an area, and how often
- Recruiting researchers - build author lists per topic and institution
- Trend analysis - count papers per category and month to measure where attention is moving
- RAG and dataset building - abstracts plus PDF links make a clean ingestion source for retrieval systems
Why arXiv
In machine learning, arXiv is effectively the publication venue: results appear there first and are cited from there. Because submissions are timestamped and revisions are tracked separately, publishedDate and updatedDate together let you distinguish genuinely new work from a revised older paper, which keyword search alone cannot.
Notes
idcarries the version suffix, for examplev1orv3, so revisions stay distinguishable.- Full abstracts are included; full paper text is not, but
pdfUrlpoints straight at it. - Papers are often cross-listed, so a single paper can appear under several
categories. - The arXiv API asks for modest request rates; the actor paces itself accordingly, so very large runs take time rather than failing.