arXiv AI & ML Research Papers Scraper — Preprints & RAG avatar

arXiv AI & ML Research Papers Scraper — Preprints & RAG

Pricing

from $1.50 / 1,000 results

Go to Apify Store
arXiv AI & ML Research Papers Scraper — Preprints & RAG

arXiv AI & ML Research Papers Scraper — Preprints & RAG

Scrape latest AI, ML, Computer Science, and Quantitative research papers from arXiv with full abstracts, authors, categories, and direct PDF URLs.

Pricing

from $1.50 / 1,000 results

Rating

0.0

(0)

Developer

Axery

Axery

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

arXiv AI & ML Research Papers Scraper

Scrape the latest preprint research papers from arXiv.org across AI, Machine Learning, NLP, Computer Vision, Robotics, and Quantitative Finance — designed for RAG pipelines, vector embedding databases, AI newsletters, and academic knowledge retrieval.

Pure HTTP via official arXiv Atom API. No browser overhead. Fast, lightweight, and reliable.


⚡ What makes this different for AI builders

  • Direct Vector Embedding Ready: Complete abstracts cleaned from linebreaks, ready to pass into embedding models (text-embedding-3, bge-large, etc.).
  • Direct PDF Links: Direct canonical PDF download links (https://arxiv.org/pdf/...).
  • Comprehensive Metadata: Extracts primaryCategory, allCategories, authors, doi, journalRef, publishedDate, and updatedDate.
  • Zero Captchas: Uses arXiv's official query protocol with automatic retry logic and backoff.

📥 Input Parameters

{
"category": "cs.AI",
"query": "large language models",
"sortBy": "submittedDate",
"maxItems": 50
}
FieldTypeDefaultDescription
categorystringcs.AIPrimary subject: cs.AI, cs.LG, cs.CL (NLP/LLMs), cs.CV, stat.ML, q-fin.ST, etc.
querystringnullOptional keyword search across titles and abstracts.
sortBystringsubmittedDatesubmittedDate (Newest first), lastUpdatedDate, or relevance.
sortOrderstringdescendingdescending or ascending.
maxItemsinteger50Maximum number of papers to extract.

📤 Output Structure

Each paper in the dataset is a clean, typed JSON object:

{
"arxivId": "2310.12345v1",
"title": "A Survey on Large Language Model Based Autonomous Agents",
"abstract": "Autonomous agents have long been a prominent research focus...",
"authors": ["Lei Wang", "Chen Ma", "Xueyang Feng", "Zeyu Zhang"],
"primaryCategory": "cs.AI",
"allCategories": ["cs.AI", "cs.CL", "cs.MA"],
"publishedDate": "2023-10-18T17:58:32Z",
"updatedDate": "2023-10-18T17:58:32Z",
"pdfUrl": "https://arxiv.org/pdf/2310.12345v1",
"absUrl": "https://arxiv.org/abs/2310.12345v1",
"doi": null,
"comment": "38 pages, 10 figures"
}