arXiv Scraper - Research Papers & Abstracts avatar

arXiv Scraper - Research Papers & Abstracts

Pricing

from $1.00 / 1,000 results

Go to Apify Store
arXiv Scraper - Research Papers & Abstracts

arXiv Scraper - Research Papers & Abstracts

Extract research papers from arXiv by keyword, category, author or date. Returns arXiv ID, title, full abstract, all authors, subject categories, publication and revision date plus links to the abstract page and the PDF. For AI research tracking, literature reviews and competitive R&D intelligence.

Pricing

from $1.00 / 1,000 results

Rating

0.0

(0)

Developer

Ryan Zinburg

Ryan Zinburg

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

0

Monthly active users

21 hours ago

Last modified

Share

arXiv Scraper - Research Papers, Abstracts & Authors

Extract paper metadata from arXiv, the preprint server where most machine learning, physics and mathematics research appears first, typically months before journal publication. Search by keyword, category, author or date window and export a complete abstract corpus.

Official arXiv API, no key required, no proxy needed.

What you get per paper

FieldExample
id2609.01234v1
titleScaling Laws for Sparse Mixture-of-Experts Models
abstractthe full abstract text
authorsevery author, in order
categories["cs.LG", "cs.CL", "stat.ML"]
publishedDatefirst submission date
updatedDatelatest revision date
arxivUrlabstract page
pdfUrldirect PDF link

Search filters

  • searchQuery - free text across title and abstract, e.g. large language models, diffusion
  • categories - arXiv categories, e.g. cs.LG, cs.CL, cs.CV, stat.ML, math.OC
  • dateFrom - earliest submission date
  • sortBy - relevance, newest submissions or latest updates
  • maxResults - how many papers to save

Example input

{
"searchQuery": "large language models",
"categories": "cs.LG,cs.CL",
"sortBy": "submittedDate",
"maxResults": 200
}

Use cases

  • AI research tracking - follow a subfield week by week without reading every listing by hand
  • Literature reviews - export an abstract corpus and search or summarise it offline
  • Competitive R&D intelligence - see which labs and authors publish in an area, and how often
  • Recruiting researchers - build author lists per topic and institution
  • Trend analysis - count papers per category and month to measure where attention is moving
  • RAG and dataset building - abstracts plus PDF links make a clean ingestion source for retrieval systems

Why arXiv

In machine learning, arXiv is effectively the publication venue: results appear there first and are cited from there. Because submissions are timestamped and revisions are tracked separately, publishedDate and updatedDate together let you distinguish genuinely new work from a revised older paper, which keyword search alone cannot.

Notes

  • id carries the version suffix, for example v1 or v3, so revisions stay distinguishable.
  • Full abstracts are included; full paper text is not, but pdfUrl points straight at it.
  • Papers are often cross-listed, so a single paper can appear under several categories.
  • The arXiv API asks for modest request rates; the actor paces itself accordingly, so very large runs take time rather than failing.