arXiv Research Paper Search Scraper
Pricing
from $1.99 / 1,000 search results
arXiv Research Paper Search Scraper
Search arXiv's public Atom API with field and sort controls, bounded pagination, stable paper normalization, safe retries, and optional proxies.
Pricing
from $1.99 / 1,000 search results
Rating
0.0
(0)
Developer
Search API
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
8 days ago
Last modified
Categories
Share
Search arXiv's public Atom API by all fields, title, author, abstract, comment, journal reference, category, report number, or arXiv ID. Sort by relevance, last update, or submission date and collect bounded, deduplicated paper records.
{"query": "large language models","searchType": "all","sortBy": "relevance","sortOrder": "descending","maxItems": 50,"pageSize": 50,"maxPages": 1,"requestDelayMs": 3000}
Each dataset row contains arXiv identity/version, title, authors, abstract and categories when present, publication/update timestamps, public abstract/paper/PDF URLs, optional DOI/journal/comment/report metadata, search/page/rank context, and scrapedAt. The fixed OUTPUT key reports run status, record/page counts, total available results, bounded failure categories, connection mode, and completion time. No-results and failures produce no placeholder paper rows.
The Actor validates response status and Atom/XML content type, enforces a 10 MB response bound, rejects malformed or changed feeds, respects hard result/page controls and an arXiv-friendly page delay, retries temporary failures with bounded backoff, and supports direct or authorized proxy access. It never stores raw blocked/error pages or logs credentials, cookies, authorization headers, proxy URLs, tokens, or sensitive response bodies. It does not bypass authentication, CAPTCHA, paywalls, regional restrictions, or other access controls.