# arXiv Research Papers & Abstracts Scraper (`scrapers_lat/arxiv-papers-scraper`) Actor

Scrape arXiv preprints by keyword, author or subject with arXiv ID, title, authors, abstract, subject categories, DOI, publication and update dates and PDF links. Export to JSON, CSV or Excel.

- **URL**: https://apify.com/scrapers\_lat/arxiv-papers-scraper.md
- **Developed by:** [Scrapers Lat](https://apify.com/scrapers_lat) (community)
- **Categories:** Developer tools, News, AI
- **Stats:** 2 total users, 1 monthly users, 75.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.55 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are a software tools running on the Apify platform, for all kinds of web data extraction and automation use cases.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

In JavaScript/TypeScript projects, use official [JavaScript/TypeScript client](https://docs.apify.com/api/client/js/docs.md):

```bash
npm install apify-client
```

In Python projects, use official [Python client library](https://docs.apify.com/api/client/python/docs.md):

```bash
pip install apify-client
```

In shell scripts, use [Apify CLI](https://docs.apify.com/cli/docs.md):

````bash
# MacOS / Linux
curl -fsSL https://apify.com/install-cli.sh | bash
# Windows
irm https://apify.com/install-cli.ps1 | iex
```bash

In AI frameworks, you might use the [Apify MCP server](https://docs.apify.com/integrations/mcp.md).

If your project is in a different language, use the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).


# README

## arXiv Research Papers & Abstracts Scraper

> Search arXiv and export clean, structured paper metadata: arXiv ID, title, authors, full abstract, subject categories, DOI, publication and update dates and PDF links. Perfect for literature reviews, research dashboards and citation tracking.

**📥 [Input](https://apify.com/scrapers_lat/arxiv-papers-scraper/input-schema) · 📤 [Output](https://apify.com/scrapers_lat/arxiv-papers-scraper/output-schema) · 💰 [Pricing](https://apify.com/scrapers_lat/arxiv-papers-scraper/pricing) · ▶️ [Examples](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples)**


![Apify](https://img.shields.io/badge/Platform-Apify-1CE1CE?logo=apify&logoColor=white)
![Research papers](https://img.shields.io/badge/Data-Research%20papers-blue)
![Output](https://img.shields.io/badge/Output-JSON%20%7C%20CSV%20%7C%20Excel-orange)

<table><tr>
<td align="center"><strong>Full abstracts</strong><br>plus authors</td>
<td align="center"><strong>Categories & DOI</strong><br>with PDF links</td>
<td align="center"><strong>JSON / CSV / Excel</strong><br>output formats</td>
</tr></table>

<br>

### What you get

Each record is one arXiv paper, ready for a reference manager, a dashboard or a spreadsheet:

- **arxivId**: the arXiv identifier (with version)
- **title**: the paper title
- **authors** and **authorCount**: the full author list
- **summary**: the complete abstract
- **categories** and **primaryCategory**: arXiv subject classifications
- **published** and **updated**: submission and last revision dates
- **doi**: the DOI when the paper has one
- **journalRef**: journal reference when available
- **comment**: author-supplied notes (pages, figures, conference)
- **pdfUrl** and **absUrl**: direct links to the PDF and the abstract page
- **observedAt**: when the record was collected

### Who is it for

| Use case | Who benefits |
|---|---|
| Literature reviews | Researchers building a corpus on a topic fast |
| Research dashboards | Teams tracking new work in a field |
| Citation and trend analysis | Analysts studying authors, categories and volume over time |
| ML datasets | Builders assembling training or evaluation sets of abstracts |

### How to use it

1. Enter a **search query** (for example `large language models`, `quantum computing` or `protein folding`).
2. Optionally choose which **field** to match (title, abstract, author or category) and how to **sort** (relevance, newest, recently updated).
3. Set **Max Items** and run. Export as JSON, CSV or Excel, or pull it through the Apify API.

### Frequently Asked Questions

**Can I search by author or subject category?**
Yes. Set the search field to Author or Category code, or pass a native query such as `au:hinton` or `cat:cs.CL`. You can also combine terms with AND / OR.

**Do I get the full abstract?**
Yes. Each record includes the complete abstract text, not just a snippet.

**Does every paper have a DOI?**
No. A DOI is included whenever the paper has one registered; otherwise the field is null. Every record still has a stable arXiv ID and links.

**How many papers can I collect?**
Set Max Items to whatever you need. Results are gathered page by page until that limit or the end of the matches is reached.

<!-- example-tasks -->
### Example use cases

Ready-to-run example tasks, each preconfigured for a common scenario. Open one and press run, or use it as a template:

- [Large Language Models research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-llm-papers-relevance): Search arXiv for large language models research papers with titles, authors, abstracts, categories, dates, and PDF links.
- [Diffusion Models research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-diffusion-models-newest): Search arXiv for diffusion models research papers with titles, authors, abstracts, categories, dates, and PDF download links.
- [Reinforcement Learning research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-reinforcement-learning-updated): Search arXiv for reinforcement learning research papers with titles, authors, abstracts, categories, dates, and PDF links.
- [Graph Neural Networks research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-graph-neural-networks-newest): Search arXiv for graph neural networks research papers with titles, authors, abstracts, categories, dates, and PDF links.
- [Quantum Computing research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-quantum-computing-relevance): Search arXiv for quantum computing research papers with titles, authors, abstracts, categories, publication dates, and PDF links.
- [Protein Folding research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-protein-folding-newest): Search arXiv for protein folding research papers with titles, authors, abstracts, categories, publication dates, and PDF links.
- [Computer Vision Transformers research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-vision-transformers-relevance): Search arXiv for computer vision transformer papers with titles, authors, abstracts, categories, dates, and PDF download links.
- [Federated Learning research papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-federated-learning-updated): Search arXiv for federated learning research papers with titles, authors, abstracts, categories, dates, and PDF download links.
- [Computation and Language papers from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-cs-cl-category-newest): Search arXiv computation and language papers in cs.CL with titles, authors, abstracts, categories, dates, and PDF links.
- [Research papers by Yoshua Bengio from arXiv](https://apify.com/scrapers_lat/arxiv-papers-scraper/examples/arxiv-author-yoshua-bengio-newest): Search arXiv for papers by Yoshua Bengio with titles, co-authors, abstracts, categories, publication dates, and PDF links.

<!-- /example-tasks -->

<!-- related-actors -->
### Related scrapers

- [Crossref Works Scraper](https://apify.com/scrapers_lat/crossref-scraper)
- [Google News Scraper](https://apify.com/scrapers_lat/google-news-scraper)

<!-- /related-actors -->

<!-- scrapers-lat-cta -->
### More scrapers at scrapers.lat

This actor is built and maintained by [scrapers.lat](https://scrapers.lat), where we publish scrapers for public platforms: finance, news, real estate, jobs, e-commerce and government data. Browse the full catalog or ask us for a custom scraper at [scrapers.lat](https://scrapers.lat).

---

> This actor is an independent tool and has no affiliation with arXiv or Cornell University. It only accesses publicly available paper metadata. Use the results in accordance with the source's terms.

# Actor input Schema

## `maxPapers` (type: `integer`):

Maximum number of papers to collect. Optional.
## `searchQuery` (type: `string`):

Words or phrase to search across arXiv papers (for example 'large language models', 'quantum computing', 'protein folding'). Advanced users can pass a native arXiv query with field prefixes such as 'ti:transformer AND cat:cs.CL'.
## `field` (type: `string`):

Which field to match the search query against when a plain phrase is entered.
## `sortBy` (type: `string`):

Order the results by relevance, newest submitted, or most recently updated.

## Actor input object example

```json
{
  "maxPapers": 50,
  "searchQuery": "large language models",
  "field": "all",
  "sortBy": "relevance"
}
````

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxPapers": 50,
    "searchQuery": "large language models"
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapers_lat/arxiv-papers-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "maxPapers": 50,
    "searchQuery": "large language models",
}

# Run the Actor and wait for it to finish
run = client.actor("scrapers_lat/arxiv-papers-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxPapers": 50,
  "searchQuery": "large language models"
}' |
apify call scrapers_lat/arxiv-papers-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=scrapers_lat/arxiv-papers-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

```json
{
    "openapi": "3.0.1",
    "info": {
        "title": "arXiv Research Papers & Abstracts Scraper",
        "description": "Scrape arXiv preprints by keyword, author or subject with arXiv ID, title, authors, abstract, subject categories, DOI, publication and update dates and PDF links. Export to JSON, CSV or Excel.",
        "version": "0.1",
        "x-build-id": "rwmA06rSYLpxAgJ5o"
    },
    "servers": [
        {
            "url": "https://api.apify.com/v2"
        }
    ],
    "paths": {
        "/acts/scrapers_lat~arxiv-papers-scraper/run-sync-get-dataset-items": {
            "post": {
                "operationId": "run-sync-get-dataset-items-scrapers_lat-arxiv-papers-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for its completion, and returns Actor's dataset items in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        },
        "/acts/scrapers_lat~arxiv-papers-scraper/runs": {
            "post": {
                "operationId": "runs-sync-scrapers_lat-arxiv-papers-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor and returns information about the initiated run in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK",
                        "content": {
                            "application/json": {
                                "schema": {
                                    "$ref": "#/components/schemas/runsResponseSchema"
                                }
                            }
                        }
                    }
                }
            }
        },
        "/acts/scrapers_lat~arxiv-papers-scraper/run-sync": {
            "post": {
                "operationId": "run-sync-scrapers_lat-arxiv-papers-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for completion, and returns the OUTPUT from Key-value store in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        }
    },
    "components": {
        "schemas": {
            "inputSchema": {
                "type": "object",
                "properties": {
                    "maxPapers": {
                        "title": "Max Papers",
                        "minimum": 1,
                        "maximum": 1000000,
                        "type": "integer",
                        "description": "Maximum number of papers to collect. Optional."
                    },
                    "searchQuery": {
                        "title": "Search query",
                        "type": "string",
                        "description": "Words or phrase to search across arXiv papers (for example 'large language models', 'quantum computing', 'protein folding'). Advanced users can pass a native arXiv query with field prefixes such as 'ti:transformer AND cat:cs.CL'."
                    },
                    "field": {
                        "title": "Search in",
                        "enum": [
                            "all",
                            "ti",
                            "abs",
                            "au",
                            "cat"
                        ],
                        "type": "string",
                        "description": "Which field to match the search query against when a plain phrase is entered.",
                        "default": "all"
                    },
                    "sortBy": {
                        "title": "Sort by",
                        "enum": [
                            "relevance",
                            "newest",
                            "lastUpdated"
                        ],
                        "type": "string",
                        "description": "Order the results by relevance, newest submitted, or most recently updated.",
                        "default": "relevance"
                    }
                }
            },
            "runsResponseSchema": {
                "type": "object",
                "properties": {
                    "data": {
                        "type": "object",
                        "properties": {
                            "id": {
                                "type": "string"
                            },
                            "actId": {
                                "type": "string"
                            },
                            "userId": {
                                "type": "string"
                            },
                            "startedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "finishedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "status": {
                                "type": "string",
                                "example": "READY"
                            },
                            "meta": {
                                "type": "object",
                                "properties": {
                                    "origin": {
                                        "type": "string",
                                        "example": "API"
                                    },
                                    "userAgent": {
                                        "type": "string"
                                    }
                                }
                            },
                            "stats": {
                                "type": "object",
                                "properties": {
                                    "inputBodyLen": {
                                        "type": "integer",
                                        "example": 2000
                                    },
                                    "rebootCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "restartCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "resurrectCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "computeUnits": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "options": {
                                "type": "object",
                                "properties": {
                                    "build": {
                                        "type": "string",
                                        "example": "latest"
                                    },
                                    "timeoutSecs": {
                                        "type": "integer",
                                        "example": 300
                                    },
                                    "memoryMbytes": {
                                        "type": "integer",
                                        "example": 1024
                                    },
                                    "diskMbytes": {
                                        "type": "integer",
                                        "example": 2048
                                    }
                                }
                            },
                            "buildId": {
                                "type": "string"
                            },
                            "defaultKeyValueStoreId": {
                                "type": "string"
                            },
                            "defaultDatasetId": {
                                "type": "string"
                            },
                            "defaultRequestQueueId": {
                                "type": "string"
                            },
                            "buildNumber": {
                                "type": "string",
                                "example": "1.0.0"
                            },
                            "containerUrl": {
                                "type": "string"
                            },
                            "usage": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "integer",
                                        "example": 1
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "usageTotalUsd": {
                                "type": "number",
                                "example": 0.00005
                            },
                            "usageUsd": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "number",
                                        "example": 0.00005
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            }
                        }
                    }
                }
            }
        }
    }
}
```
