arXiv Scraper - Papers, Authors, Abstracts & PDFs avatar

arXiv Scraper - Papers, Authors, Abstracts & PDFs

Pricing

from $1.10 / 1,000 paper scrapeds

Go to Apify Store
arXiv Scraper - Papers, Authors, Abstracts & PDFs

arXiv Scraper - Papers, Authors, Abstracts & PDFs

Scrape arXiv papers by search query or ID from the official API: title, abstract, authors, categories, publication date, DOI, journal reference and PDF link. No key, no browser, cents per paper.

Pricing

from $1.10 / 1,000 paper scrapeds

Rating

0.0

(0)

Developer

Scrape Sage

Scrape Sage

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Share

arXiv Scraper - Papers, Authors, Abstracts & PDFs

Search arXiv and get clean, structured rows for every paper: title, full abstract, authors, categories, publication date, DOI and journal reference where available, and a direct PDF link. Search by keyword (or advanced field syntax), or fetch exact papers by arXiv ID. Built on the official arXiv API - no key, no browser.

What you get per paper

FieldMeaning
title / summaryTitle and the full abstract
authors / authorCountAuthor list and count
primaryCategory / categoriese.g. cs.AI, cs.LG, stat.ML
published / updatedFirst submission and last revision dates
pdfUrl / arxivUrlDirect PDF and the abstract page
doi / journalRef / commentDOI and journal reference once published; author comment

Input

{ "searchQueries": ["large language models"], "maxItemsPerQuery": 100, "sortBy": "submitted" }
  • Search queries - one per line. Plain terms work, and so does arXiv field syntax: cat:cs.AI, au:hinton, ti:transformer, abs:diffusion.
  • arXiv IDs - or fetch specific papers by ID (2312.00752).
  • Import from a file - paste a whole list, or link a public .txt/.csv, a Google Sheet/Drive link, or an Apify key-value-store record.
  • Sort by relevance, newest submitted, or recently updated. Max papers per query / total bound the run.
  • Output fields - tick only the columns you want.

Leave everything empty and the run returns a small free sample so you can see the shape first.

Reliability

Reads the official arXiv API - a public Atom feed, no key, no proxy, no anti-bot. The actor paces its requests to stay within arXiv's guidance. A run that returns nothing bills $0.

Honest limits

  • doi and journalRef are only present once a preprint is formally published in a journal - most arXiv preprints have neither, which is the true state of the record, not a gap.
  • Search covers arXiv's index, not the full text of the PDFs; you get the abstract and metadata plus the PDF link to download separately.

Pricing

$0.002 per paper on the FREE tier (tiered pricing lowers it with volume). Only papers actually saved are billed; an empty run costs nothing.

Output views

  • Papers - title, authors, category, date, DOI and the PDF/abstract links.

Use with AI assistants (MCP)

Available through the Apify MCP server - an agent can pull the latest papers on a topic (with abstracts and PDF links) for a literature review or RAG pipeline.

Agent-ready: autonomous payments (x402 & Skyfire)

This actor is agent-ready - AI agents can discover it, run it, and pay for it autonomously, with no Apify account and no human in the loop. It uses pay-per-event pricing and limited permissions, so it qualifies for Apify's agentic-payment standards:

  • x402 - an open, HTTP-native payment protocol. Agents pay per run in USDC on the Base network directly through the Apify MCP server - no account, no API key.
  • Skyfire - agent-to-service payments for fully autonomous AI-agent workflows.

Building an AI agent, MCP tool, or autonomous data pipeline? This scraper is ready to plug in and pay as it goes.