arXiv Research Paper Search Scraper avatar

arXiv Research Paper Search Scraper

Pricing

from $1.99 / 1,000 search results

Go to Apify Store
arXiv Research Paper Search Scraper

arXiv Research Paper Search Scraper

Search arXiv's public Atom API with field and sort controls, bounded pagination, stable paper normalization, safe retries, and optional proxies.

Pricing

from $1.99 / 1,000 search results

Rating

0.0

(0)

Developer

Search API

Search API

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

8 days ago

Last modified

Categories

Share

Search arXiv's public Atom API by all fields, title, author, abstract, comment, journal reference, category, report number, or arXiv ID. Sort by relevance, last update, or submission date and collect bounded, deduplicated paper records.

{
"query": "large language models",
"searchType": "all",
"sortBy": "relevance",
"sortOrder": "descending",
"maxItems": 50,
"pageSize": 50,
"maxPages": 1,
"requestDelayMs": 3000
}

Each dataset row contains arXiv identity/version, title, authors, abstract and categories when present, publication/update timestamps, public abstract/paper/PDF URLs, optional DOI/journal/comment/report metadata, search/page/rank context, and scrapedAt. The fixed OUTPUT key reports run status, record/page counts, total available results, bounded failure categories, connection mode, and completion time. No-results and failures produce no placeholder paper rows.

The Actor validates response status and Atom/XML content type, enforces a 10 MB response bound, rejects malformed or changed feeds, respects hard result/page controls and an arXiv-friendly page delay, retries temporary failures with bounded backoff, and supports direct or authorized proxy access. It never stores raw blocked/error pages or logs credentials, cookies, authorization headers, proxy URLs, tokens, or sensitive response bodies. It does not bypass authentication, CAPTCHA, paywalls, regional restrictions, or other access controls.