# arXiv Papers Scraper – Search Research Papers & Abstracts (`stevenkramp/arxiv-papers-scraper`) Actor

Search arXiv research papers by keywords, categories (cs.AI, cs.CL, stat.ML …) and submission date, or fetch papers by ID: title, abstract, authors, categories, dates, PDF link, DOI and journal reference. Official arXiv API, CC0 metadata. Pay only per paper returned.

- **URL**: https://apify.com/stevenkramp/arxiv-papers-scraper.md
- **Developed by:** [Steven Kramp](https://apify.com/stevenkramp) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 papers

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## arXiv Papers Scraper – Search Research Papers & Abstracts

Search **arXiv** research papers by keywords, categories and submission date, or fetch specific papers by ID. You get clean, structured metadata with the full abstract, authors, categories, dates, PDF link, DOI and journal reference. The data comes from **arXiv's official API**, and arXiv metadata is free to reuse (CC0).

**$2 per 1,000 papers.** Up to 10,000 papers per run. You only pay for papers returned.

### What you get for each paper

| Field | Example |
|---|---|
| `arxivId`, `version`, `absUrl`, `pdfUrl` | 1706.03762 · v7 |
| `title`, `abstract` | Attention Is All You Need · full abstract |
| `authors` | list of author names |
| `primaryCategory`, `categories` | cs.CL · cs.CL, cs.LG |
| `published`, `updated` | submission and last update date |
| `doi`, `journalRef`, `comment` | when the authors provided them |
| `query`, `position`, `scrapedAt` | run context |

### Use cases

- **Research monitoring:** get every new paper in your field each day or week with a schedule ("cs.AI", "llm agents").
- **Literature reviews:** pull hundreds of papers with abstracts into a spreadsheet or reference manager.
- **AI and RAG pipelines:** feed fresh abstracts and PDF links into your own models and knowledge bases.
- **Trend analysis:** count papers per topic and month to see which research areas grow.

### How to use

```json
{
  "searchQuery": "llm agents",
  "categories": ["cs.AI", "cs.CL"],
  "dateFrom": "2026-01-01",
  "sortBy": "newest",
  "maxResults": 500
}
```

Advanced arXiv syntax works too, e.g. `ti:transformer AND au:vaswani`. To get specific papers, list IDs or URLs in `arxivIds`.

### Pricing

Pay per event: **$0.002 per paper returned** ($2 per 1,000). Platform usage is included.

### Use with AI agents (MCP)

This Actor works as a tool for AI assistants and agents – Claude, ChatGPT, Cursor, VS Code, n8n and other MCP clients – through Apify's hosted MCP server. Add this server URL to your client:

```
https://mcp.apify.com?tools=stevenkramp/arxiv-papers-scraper
```

Sign in with your Apify account when asked. Your agent can then call the Actor in plain language, for example: *"Find this week's new arXiv papers on AI agents and summarize the five most interesting abstracts."* Runs started by your agent are normal Actor runs on your Apify account at the same pay-per-event price.

### More from stevenkramp

Other Actors by the same developer – same quality standards, pay only for results:

**Search & trends**

- [Google Trends Scraper](https://apify.com/stevenkramp/google-trends-scraper) – interest over time, regions, rising queries
- [Keyword Trends Finder](https://apify.com/stevenkramp/keyword-trends-finder) – keyword ideas with trend direction
- [Google News Scraper](https://apify.com/stevenkramp/google-news-scraper) – news articles with real URLs
- [Google Images Scraper](https://apify.com/stevenkramp/google-images-scraper) – full-size image URLs
- [Google Shopping Scraper](https://apify.com/stevenkramp/google-shopping-scraper) – prices and merchants
- [Google Jobs Scraper](https://apify.com/stevenkramp/google-jobs-scraper) – job listings

**Apps**

- [Google Play Store Scraper](https://apify.com/stevenkramp/google-play-store-scraper) – Android app data and rankings
- [Google Play Reviews Scraper](https://apify.com/stevenkramp/google-play-reviews-scraper) – Play Store reviews and ratings
- [Apple App Store Scraper](https://apify.com/stevenkramp/apple-app-store-scraper) – iPhone, iPad and Mac app data
- [Shopify App Store Scraper](https://apify.com/stevenkramp/shopify-app-store-scraper) – Shopify apps and pricing plans

**Research & media**

- [Apple Podcasts Scraper](https://apify.com/stevenkramp/apple-podcasts-scraper) – podcasts with latest episodes

**Websites & places**

- [Website SEO Audit](https://apify.com/stevenkramp/website-seo-audit) – broken links, titles, sitemap
- [Germany Neighborhood Profile](https://apify.com/stevenkramp/germany-neighborhood-profile) – German neighborhood rents and vacancy

### FAQ

**How fast is it?** arXiv asks for one request every 3 seconds, and we follow this. One request returns up to 200 papers, so 1,000 papers take about 15 seconds.

**Full texts?** We return the abstract and the PDF link. The PDFs are hosted by arXiv. Please respect each paper's license when you reuse full texts.

**Something broken?** Open an issue in the Issues tab. We fix problems quickly.

Thank you to arXiv for use of its open access interoperability.

# Actor input Schema

## `searchQuery` (type: `string`):

Keywords, e.g. "llm agents" (all words must match). Advanced arXiv syntax also works, e.g. ti:transformer AND au:hinton.

## `categories` (type: `array`):

arXiv categories, e.g. cs.AI, cs.CL, cs.LG, stat.ML, q-fin.TR. Papers in any of them match.

## `dateFrom` (type: `string`):

Format YYYY-MM-DD.

## `dateTo` (type: `string`):

Format YYYY-MM-DD.

## `sortBy` (type: `string`):

Order of the results.

## `maxResults` (type: `integer`):

Up to 10,000 per run. You only pay for papers returned.

## `arxivIds` (type: `array`):

Fetch specific papers, e.g. 2610.06844 or https://arxiv.org/abs/1706.03762.

## Actor input object example

```json
{
  "searchQuery": "llm agents",
  "categories": [
    "cs.AI"
  ],
  "sortBy": "newest",
  "maxResults": 100
}
```

# Actor output Schema

## `items` (type: `string`):

One item per paper.

## `overview` (type: `string`):

Key fields as a table.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQuery": "llm agents",
    "categories": [
        "cs.AI"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("stevenkramp/arxiv-papers-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQuery": "llm agents",
    "categories": ["cs.AI"],
}

# Run the Actor and wait for it to finish
run = client.actor("stevenkramp/arxiv-papers-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQuery": "llm agents",
  "categories": [
    "cs.AI"
  ]
}' |
apify call stevenkramp/arxiv-papers-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,stevenkramp/arxiv-papers-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/1lkm9bI9pwhWTqWNg/builds/ztKyYwaxGlYpkJ3yB/openapi.json
