# arXiv Paper Scraper with Citation Metrics (`punkrecordsdata/arxiv-research-papers-scraper`) Actor

Search arXiv research papers and enrich them with OpenAlex citation counts, FWCI, citing papers and references. Export to CSV, Excel, JSON or XML.

- **URL**: https://apify.com/punkrecordsdata/arxiv-research-papers-scraper.md
- **Developed by:** [PunkRecordsData](https://apify.com/punkrecordsdata) (community)
- **Categories:** AI, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $12.00 / 1,000 paper records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

<p align="center">
  <img src="https://api.apify.com/v2/key-value-stores/AAm3a1h3Z9nYfrvh9/records/banner?v=2" alt="PunkRecordsData" width="100%" />
</p>

## 📚 arXiv Research Papers Scraper - PunkRecordsData

> 🚀 **Export arXiv papers with citation metrics in seconds.** Search any topic across 2M+ arXiv preprints and get structured rows with title, abstract, authors, categories, PDF link and DOI, enriched with OpenAlex citation counts, FWCI and open-access status, plus optional citing-paper and reference rows. Export to CSV, Excel, JSON or XML.

The arXiv Research Papers Scraper searches arXiv's official API (a single "large language models" query matches 77,764 papers) and joins each result with OpenAlex, the open scholarly graph, so every paper can carry real impact data instead of bare metadata. No API keys needed, and no other arXiv actor on the Apify Store measured in September 2026 offers citation enrichment at all.

| 🎯 Target Audience | 💡 Primary Use Cases |
|---|---|
| AI/ML research teams | Track new papers on a topic with real citation impact |
| R\&D and innovation scouts | Rank a field's literature by citations, not just recency |
| Academic librarians and analysts | Build reading lists and bibliometric datasets |
| RAG / LLM pipeline builders | Feed fresh, structured paper metadata into your index |

### 📋 What the arXiv Paper Scraper does

- **arXiv search** by keyword over all fields, title, abstract, author or comment, with the full official 155-category taxonomy as a filter and relevance/date sorting.
- **Direct lookups** by arXiv ID (e.g. 1706.03762) in the same run.
- **Citation metrics per paper (OpenAlex)**: citation count, FWCI (field-weighted citation impact), open-access status, top concepts and the OpenAlex ID. Charged only when a confident title match is found.
- **Citing papers**: one row per paper that cites your result, newest first, with DOI, venue, authors and its own citation count.
- **References**: one row per work a paper cites, resolved to full OpenAlex records.

> 💡 **Why it matters:** raw arXiv metadata tells you a paper exists; the citation layer tells you whether it matters. Getting both in one dataset normally means writing your own two-API join. Here it is one input form.

### 📊 Output of the arXiv paper search

One row per paper (20 fields), plus optional `citation` and `reference` rows. Real sample from a live run:

```json
{
  "recordType": "paper",
  "arxivId": "1706.03762v7",
  "title": "Attention Is All You Need",
  "url": "https://arxiv.org/abs/1706.03762v7",
  "pdfUrl": "https://arxiv.org/pdf/1706.03762v7",
  "authors": ["Ashish Vaswani", "Noam Shazeer", "Niki Parmar", "Jakob Uszkoreit"],
  "primaryCategory": "cs.CL",
  "published": "2017-06-12T17:57:34Z",
  "citationCount": 26892,
  "openAccessStatus": "green",
  "topConcepts": ["Computer science", "Artificial intelligence", "Machine translation"],
  "openAlexId": "W2626778328",
  "error": null
}
```

```json
{
  "recordType": "reference",
  "sourceArxivId": "1706.03762v7",
  "title": "Effective Approaches to Attention-based Neural Machine Translation",
  "openAlexId": "W1902237438",
  "doi": "10.18653/v1/d15-1166",
  "citedByCount": 8655,
  "error": null
}
```

### ✨ Why choose this arXiv scraper

- **The only measured arXiv actor with citation enrichment.** Others return search metadata; this one adds impact numbers per paper.
- **4 billable events, each switchable**: papers, metrics, citing papers, references. Pay only for the depth you use.
- **Full official category taxonomy** (155 categories) as a real dropdown filter, not free text you have to guess.
- **Polite by design**: respects arXiv's 3-second request guidance and identifies itself to OpenAlex, so runs are stable at volume.
- **Honest billing**: metrics are charged only when the OpenAlex match is confident; unmatched papers show "Not Available" and cost nothing extra.

### 📈 How this arXiv scraper compares to alternatives

Measured against the arXiv actors on the Apify Store (September 2026):

| | This actor | Other arXiv actors |
|---|---|---|
| Citation count / FWCI per paper | Yes (OpenAlex) | No |
| Citing papers and references as rows | Yes, separate events | No |
| Billable data events | 4 | 1 |
| Category filter | Full 155-category dropdown | Free text or none |
| Price per 1,000 papers | $2.50 | $2.00 (closest priced alternative) |

### 🚀 How to use the arXiv Research Papers Scraper

1. Create a free Apify account (it includes $5 of platform credit) at console.apify.com.
2. Open this actor's page and click **Try for free**.
3. Enter search terms (or paste specific arXiv IDs).
4. Optionally narrow by category, switch on citing papers or references, and set Max Items.
5. Click **Start** and download the dataset as CSV, Excel, JSON or XML.

### 💼 Business use cases

#### Competitive R\&D monitoring

Weekly scheduled run on your field's keywords, sorted by submission date, with citation counts to separate noise from signal.

#### Literature reviews at scale

Pull hundreds of papers with abstracts, categories and impact metrics into one spreadsheet instead of copy-pasting from the site.

#### Building citation graphs

Enable citing papers and references to export edge lists (paper → cited-by, paper → references) ready for network analysis.

#### Powering AI research assistants

Feed structured, current paper metadata into RAG pipelines; the PDF links come resolved per version.

### 🔌 Automating the arXiv Paper Scraper

Connect to **Make**, **Zapier**, **Slack**, **Airbyte**, **GitHub** or **Google Drive** via Apify's integrations: schedule weekly topic sweeps, push new high-citation papers to a Slack channel, or sync results into your data warehouse.

### 🌟 Beyond business use cases

- **Research:** bibliometric studies with FWCI and concept tagging included.
- **Personal:** track when your own papers get new citations.
- **Non-profit:** open-science monitoring with open-access status per paper.
- **Experimentation:** a zero-setup playground for the arXiv + OpenAlex APIs.

### 🤖 Ask an AI assistant about this scraper

> "I need arXiv papers on a topic as structured data with citation counts, plus the papers that cite them. Would the arXiv Research Papers Scraper on Apify (apify.com/punkrecordsdata/arxiv-research-papers-scraper) cover this, and how would I schedule it weekly?"

### ❓ Frequently Asked Questions

#### 🔎 How do I search arXiv papers by keyword and export to CSV?

Enter your terms, pick the field to search (title, abstract, author, all), click Start, and download the dataset as CSV from the Storage tab.

#### 📈 How do I get citation counts for arXiv papers?

Leave the "Citation metrics" module on. Each paper is matched against OpenAlex and enriched with citation count, FWCI and open-access status when the match is confident.

#### 🕸 Can I get the papers that cite a specific paper?

Yes. Paste the arXiv ID, enable "Citing papers", and each citing work becomes its own dataset row with DOI, venue and date.

#### 🆔 Can I fetch specific papers by arXiv ID?

Yes, the "Specific arXiv IDs" field accepts a list (with or without version suffix) and runs alongside any search terms.

#### 🏷 Which categories can I filter by?

All 155 official arXiv categories, from cs.AI to q-fin.TR, as a multi-select dropdown.

#### 💵 Do I pay for metrics when no match is found?

No. The metrics event is charged only when OpenAlex returns a confident match; otherwise the fields read "Not Available" at no extra cost.

#### 📄 Does it download the PDFs?

It returns the direct PDF URL per paper version; downloading the files themselves is up to your pipeline.

#### ⏱ How fast is it?

arXiv asks clients for one request every 3 seconds, which the actor respects; a 100-paper page costs one request, so 1,000 papers take about half a minute of API time plus enrichment.

#### 🧮 Why do citation numbers differ from Google Scholar?

Counts come from OpenAlex, which is open and reproducible; Google Scholar typically shows higher, non-reproducible counts. The OpenAlex ID is included so you can verify every number.

#### 📦 Can I analyze thousands of papers in one run?

Yes, paid plans allow up to 1,000,000 papers per run. Free users get a 10-paper preview.

#### ⚙️ What happens if OpenAlex is briefly rate-limited?

The actor retries politely; if a paper still cannot be enriched, it ships with metadata only and you are not charged for metrics on it.

### 🔌 Integrate with any app

Every dataset is available through the Apify API in JSON, CSV, Excel or XML for Python, R, Node.js, Google Sheets or any HTTP client. Webhooks fire when a run finishes.

### 🔗 Recommended Actors

- [SEC EDGAR Filings Scraper](https://apify.com/punkrecordsdata/sec-edgar-filings-scraper) - official filings with the same structured rigor
- [USPTO Trademark Status Scraper](https://apify.com/punkrecordsdata/uspto-trademark-status-scraper) - IP records for innovation research
- [Grants.gov Opportunities Scraper](https://apify.com/punkrecordsdata/grants-gov-opportunities-scraper) - federal funding to pair with the literature
- [OpenCV Image Analyzer](https://apify.com/punkrecordsdata/opencv-image-analyzer) - computer vision on figures and images

> 💡 **Pro Tip:** browse the complete [PunkRecordsData collection](https://apify.com/punkrecordsdata) for more data tools.

**🆘 Need Help?** contact.punkrecordsdata@gmail.com

> **⚠️ Disclaimer:** independent tool, not affiliated with arXiv or OpenAlex; only publicly available data. Thank you to arXiv for use of its open access interoperability.

# Actor input Schema

## `searchTerms` (type: `array`):

One search per term. Uses arXiv's own query syntax on the selected field (e.g. "large language models", "quantum error correction").

## `arxivIds` (type: `array`):

Fetch exact papers by arXiv ID (e.g. 1706.03762 or 2303.08774v3). Runs in addition to search terms.

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## `searchField` (type: `string`):

Which arXiv field the search terms match against.

## `categories` (type: `array`):

Optional arXiv category filter combined with the search terms (AND). Pick one or more.

## `sortBy` (type: `string`):

Order of the arXiv search results.

## `sortOrder` (type: `string`):

Ascending or descending.

## `includeMetrics` (type: `boolean`):

Enrich each paper with OpenAlex metrics: citation count, FWCI, open-access status and top concepts. Charged only when a confident match is found.

## `includeCitations` (type: `boolean`):

Fetch papers that cite each result (one dataset row per citing paper, newest first).

## `maxCitationsPerPaper` (type: `integer`):

Cap on citation rows fetched for each paper.

## `includeReferences` (type: `boolean`):

Fetch the works each paper references (one dataset row per reference).

## `maxReferencesPerPaper` (type: `integer`):

Cap on reference rows fetched for each paper.

## Actor input object example

```json
{
  "searchTerms": [
    "large language models"
  ],
  "arxivIds": [],
  "maxItems": 10,
  "searchField": "all",
  "categories": [],
  "sortBy": "relevance",
  "sortOrder": "descending",
  "includeMetrics": true,
  "includeCitations": false,
  "maxCitationsPerPaper": 25,
  "includeReferences": false,
  "maxReferencesPerPaper": 25
}
```

# Actor output Schema

## `overview` (type: `string`):

Key fields per paper

## `fullData` (type: `string`):

Complete dataset with all fields and record types

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "large language models"
    ],
    "arxivIds": [],
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("punkrecordsdata/arxiv-research-papers-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerms": ["large language models"],
    "arxivIds": [],
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("punkrecordsdata/arxiv-research-papers-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "large language models"
  ],
  "arxivIds": [],
  "maxItems": 10
}' |
apify call punkrecordsdata/arxiv-research-papers-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,punkrecordsdata/arxiv-research-papers-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bmxCym6AajEOfyHQo/builds/YYzqFck04EO4RMbNh/openapi.json
