ArXiv Papers Scraper
Pricing
from $0.48 / 1,000 paper extracteds
ArXiv Papers Scraper
Search arXiv by topic, author, category, date, or paper ID and export structured paper metadata for literature monitoring and research datasets.
Pricing
from $0.48 / 1,000 paper extracteds
Rating
0.0
(0)
Developer
Stas Persiianenko
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Search and export arxiv papers as structured records for literature monitoring, systematic review preparation, and research datasets.
The Actor queries the official public arXiv Atom API. Search by topic, author, category, submission date, or known paper ID, then save titles, abstracts, ordered authors, categories, dates, DOI metadata, and PDF links to an Apify dataset.
No arXiv login, cookie, API key, or proxy is required.
What can ArXiv Papers Scraper do?
- Search titles, abstracts, and author names with a topic query.
- Filter papers by one or more arXiv categories.
- Filter by author and first-submission date.
- Retrieve known papers by arXiv ID, abstract URL, or PDF URL.
- Sort by submission date, last update, or relevance.
- Export up to 500 unique papers per run.
- Normalize Atom metadata into integration-ready JSON.
- Provide canonical abstract and PDF links.
- Preserve DOI, journal reference, and submission comments when available.
- Support repeatable scheduled literature-monitoring runs.
Who is it for?
Researchers build focused reading lists without copying metadata by hand.
Librarians and research-support teams collect bounded topic or category exports for discovery workflows.
Data scientists create reproducible metadata datasets for analysis, classification, or citation-pipeline inputs.
R&D teams schedule a query and compare exports to identify newly submitted work.
Developers and AI agents retrieve clean records through the Apify API or MCP instead of parsing Atom XML.
This Actor searches public arXiv metadata. It does not submit papers, access private accounts, or bypass arXiv login.
Why use this Actor?
The source API is public, but production workflows still need query construction, pagination, validation, rate-limit handling, metadata normalization, deduplication, storage, and integrations.
This Actor packages those steps into one reusable task:
- validate the requested search scope;
- build a structured arXiv query;
- fetch bounded Atom pages politely;
- retry temporary failures;
- normalize and deduplicate papers;
- charge only for saved records;
- write typed JSON to the default dataset.
It fails clearly rather than returning fabricated or cached research records after an upstream error.
Input parameters
| Field | Type | Description |
|---|---|---|
query | string | Topic text matched across titles, abstracts, and author names. |
author | string | Author-name filter, such as Geoffrey Hinton. |
categories | string[] | arXiv category codes; multiple values are combined with OR. |
paperIds | string[] | Known IDs or arXiv abs/PDF URLs. |
dateFrom | date | Earliest first-submission date, YYYY-MM-DD. |
dateTo | date | Latest first-submission date, YYYY-MM-DD. |
sortBy | enum | submittedDate, lastUpdatedDate, or relevance. |
sortOrder | enum | descending or ascending. |
maxItems | integer | Maximum unique records, from 1 to 500. |
maxRequestRetries | integer | Temporary-request retries, from 0 to 5. |
Provide at least one of query, author, categories, or paperIds.
Search filters are combined with AND. Multiple categories are combined with OR.
When paperIds are supplied, the Actor retrieves those exact records and applies the topic, author, category, and date filters locally as well.
Quick start
- Open the Actor in Apify Console.
- Enter a topic such as
graph neural networks. - Optionally add
cs.LGto arXiv categories. - Choose a maximum number of papers.
- Click Start.
- Open the Dataset tab to preview or download JSON, CSV, Excel, XML, or RSS.
Example input:
{"query": "graph neural networks","categories": ["cs.LG"],"sortBy": "submittedDate","sortOrder": "descending","maxItems": 25}
Search by author
Use the dedicated author field instead of embedding arXiv query syntax:
{"author": "Geoffrey Hinton","sortBy": "submittedDate","sortOrder": "descending","maxItems": 20}
The filter is sent to the official author index. The returned record keeps the complete ordered author list.
Search arXiv math and science categories
Category codes are source-specific identifiers such as:
cs.AI— Artificial Intelligencecs.CL— Computation and Languagecs.LG— Machine Learningmath.OC— Optimization and Controlquant-ph— Quantum Physicsstat.ML— Machine Learning in Statistics
Example category export:
{"categories": ["math.OC", "stat.ML"],"sortBy": "submittedDate","maxItems": 100}
Consult arXiv's current category taxonomy when choosing codes. Invalid category shapes are rejected before a source request.
Monitor papers in a submission-date window
Use a bounded date range for repeatable reviews:
{"query": "retrieval augmented generation","categories": ["cs.AI"],"dateFrom": "2025-01-01","dateTo": "2025-12-31","sortBy": "submittedDate","sortOrder": "descending","maxItems": 100}
For recurring monitoring, schedule this Actor in Apify and move the date window forward. Downstream automation can compare paperId and updatedAt with an earlier dataset.
The Actor does not itself send alerts or calculate changes between runs.
Retrieve known paper IDs
Use IDs when you already have a reading list:
{"paperIds": ["1706.03762","https://arxiv.org/abs/2303.08774","https://arxiv.org/pdf/2302.13971.pdf"],"maxItems": 3}
Versioned identifiers such as 1706.03762v7 are supported.
Legacy-style arXiv identifiers are also accepted.
Output fields
| Field | Meaning |
|---|---|
paperId | arXiv ID, including returned version. |
title | Normalized paper title. |
abstract | Normalized abstract text. |
authors | Ordered array of author names. |
categories | All assigned arXiv category codes. |
primaryCategory | Primary category code. |
publishedAt | Initial submission timestamp. |
updatedAt | Latest version timestamp. |
pdfUrl | Direct PDF URL. |
absUrl | Canonical abstract-page URL. |
doi | DOI when arXiv provides one; otherwise null. |
journalReference | Journal reference when available. |
comment | Source submission comment when available. |
query | Structured query used, or null for ID retrieval. |
fetchedAt | UTC extraction timestamp. |
Example output
A real topic-search result has this shape:
{"paperId": "2608.27413v1","title": "Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling","abstract": "Friend recommendation is inherently graph-structured...","authors": ["Maksim Utushkin", "Andrei Ovsiannikov", "Alexander D'yakonov"],"categories": ["cs.IR", "cs.LG", "cs.SI"],"primaryCategory": "cs.IR","publishedAt": "2026-08-27T17:41:33Z","updatedAt": "2026-08-27T17:41:33Z","pdfUrl": "https://arxiv.org/pdf/2608.27413v1","absUrl": "https://arxiv.org/abs/2608.27413v1","doi": null,"journalReference": null,"comment": "12 pages, 4 figures, 8 tables...","query": "all:\"graph neural networks\" AND (cat:cs.LG)","fetchedAt": "2026-08-28T06:36:49.711Z"}
Optional metadata remains null when the paper's arXiv entry does not supply it.
How much does it cost to export arXiv papers?
The Actor uses pay per event:
- $0.001 when a run starts;
- $0.0008 per saved paper on the BRONZE tier.
Approximate BRONZE examples:
| Saved papers | Estimated price |
|---|---|
| 5 | $0.005 |
| 25 | $0.021 |
| 100 | $0.081 |
| 500 | $0.401 |
A no-result run pays only the start fee. Failed input validation happens before source extraction, while the live pricing system remains the final billing authority.
Higher Apify subscription tiers receive the lower item prices shown in Console.
Export to spreadsheets and data pipelines
Every result is stored in the default Apify dataset.
From the Dataset tab you can download:
- JSON for applications and archives;
- CSV or Excel for analysts;
- XML for legacy systems;
- RSS for compatible readers.
Use Apify integrations to send completed datasets to Google Drive, Slack, webhooks, Zapier, Make, or your own service.
For durable pipelines, store paperId as the stable join key and inspect updatedAt for version changes.
Schedule literature monitoring
A practical recurring workflow is:
- create a saved task with a topic, category, and date range;
- add a daily or weekly schedule;
- receive a run-finished webhook;
- fetch the dataset through the API;
- compare paper IDs with your existing catalog;
- route new papers to a review queue.
Keep date windows explicit for reproducible monitoring.
ArXiv publication timing and source API indexing determine freshness.
Run with the Apify API
cURL
curl -X POST \"https://api.apify.com/v2/acts/automation-lab~arxiv-paper-search-export/runs?token=$APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"query":"graph neural networks","categories":["cs.LG"],"maxItems":25}'
JavaScript
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('automation-lab/arxiv-paper-search-export').call({query: 'graph neural networks',categories: ['cs.LG'],maxItems: 25,});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
Python
from apify_client import ApifyClientclient = ApifyClient(token="YOUR_APIFY_TOKEN")run = client.actor("automation-lab/arxiv-paper-search-export").call(run_input={"query": "graph neural networks","categories": ["cs.LG"],"maxItems": 25,})items = client.dataset(run["defaultDatasetId"]).list_items().itemsprint(items)
Never commit an Apify token to source control.
Use through MCP
Add this Actor to Claude Code:
claude mcp add --transport http apify \"https://mcp.apify.com?tools=automation-lab/arxiv-paper-search-export"
Claude Desktop, Cursor, and VS Code setup
Use this MCP configuration in Claude Desktop, Cursor, or VS Code:
{"mcpServers": {"apify": {"url": "https://mcp.apify.com?tools=automation-lab/arxiv-paper-search-export"}}}
Example prompts:
- "Find five recent cs.LG papers about graph neural networks and summarize their metadata."
- "Export papers by Geoffrey Hinton with their arXiv PDF links."
- "Collect quant-ph papers from this date range for my literature review."
MCP invokes the same validated Actor input and returns the same dataset records.
Reliability and source etiquette
The implementation uses the official public Atom API, bounded request pages, a 30-second request timeout, and increasing retry delays.
For multi-page runs it waits at least three seconds between source requests in line with arXiv API guidance.
The Actor does not need browser rendering, residential proxies, or account sessions.
Temporary rate limits can still occur. The Actor retries only the bounded number requested and then fails with a clear error instead of silently returning stale data.
Limits
maxItemsis capped at 500 per run.- Search is governed by arXiv's index and query semantics.
- Topic text is treated as one quoted phrase for precise matching.
- Date filters use first-submission time, not the latest revision time.
- DOI, journal, and comment fields are optional source metadata.
- This Actor exports metadata and links; it does not download PDF files.
- It does not calculate citations, affiliations, full-text entities, or semantic similarity.
- It does not send alerts or merge datasets across runs.
Troubleshooting
The run says I must provide a search field
Supply at least one of query, author, categories, or paperIds. The Actor intentionally rejects unbounded blank searches.
My category is rejected
Use an arXiv category code such as cs.AI, math.OC, or quant-ph, not a descriptive category name.
A search returns no papers
Try a broader phrase, remove one filter, verify the category code, or widen the date range. All active filters are combined with AND.
The arXiv API rate-limited the run
Wait before retrying, keep retry count bounded, and avoid overlapping large scheduled runs. The Actor already applies polite pagination delays.
A paper has no DOI or journal reference
Those fields are optional in arXiv. null means the source entry did not include the value.
Legality and responsible use
ArXiv exposes public scholarly metadata for discovery and reuse, but users remain responsible for their workflow.
- Follow arXiv's API terms and rate guidance.
- Respect paper copyrights and licenses when following PDF links.
- Do not infer sensitive personal information about authors.
- Verify important bibliographic details against the source record.
- Attribute arXiv and original authors where appropriate.
This Actor is an independent automation tool and is not affiliated with or endorsed by Cornell University or arXiv.
Related Actor
For broader scholarly web search and citation-result metadata, use Google Scholar Scraper.
Choose this Actor when you need official arXiv metadata and category/date filtering. Choose Google Scholar Scraper when you need discovery across publishers and repositories.
FAQ
Does this Actor require an arXiv login?
No. It uses public metadata and never asks for account credentials.
Can I search by paper title?
Yes. Put the title or a distinctive phrase in query. Topic search covers titles and abstracts.
Can I fetch a specific version?
Yes. Supply a versioned ID such as 1706.03762v7.
Can I export PDFs?
The output includes a direct pdfUrl, but the Actor does not download or store PDF binaries.
Are multiple categories AND or OR?
Multiple category values are OR. Category filtering is then ANDed with topic, author, and date filters.
How do I detect revised papers?
Store paperId and updatedAt, then compare them across scheduled datasets. Versioned IDs may also change when arXiv returns a newer version.
Is the output suitable for a research dataset?
It is suitable as structured source metadata. Review source licenses, document your query and collection date, and validate records for your methodology.
What happens if the source is unavailable?
The Actor performs bounded retries and fails the run if it cannot obtain valid Atom data. It does not substitute fabricated or cached results.