arXiv Scraper
DeprecatedPricing
from $0.70 / 1,000 results
arXiv Scraper
DeprecatedSearch and extract academic papers from arXiv.org. Get paper titles, authors, abstracts, categories, and PDF links for AI/ML, physics, math, and more.
Pricing
from $0.70 / 1,000 results
Rating
0.0
(0)
Developer
Artificially
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
5 days ago
Last modified
Categories
Share
arXiv Papers Scraper - Enhanced
Search and extract academic papers from arXiv.org with citation analysis, author profiles, and impact metrics via Semantic Scholar integration.
Use with AI agents (MCP)
This actor works as a tool for AI agents through Apify's MCP server at https://mcp.apify.com, so Claude, ChatGPT, Cursor and other MCP clients can search arXiv papers directly. For agents, set "compactOutput": true and keep maxPapers small (5-20) to get short, token-friendly results.
Quick setup (sign in with your Apify account when asked):
- Claude (claude.ai or Claude Desktop): Settings → Connectors → Add custom connector, and paste
https://mcp.apify.com?tools=artificially/arxiv-scraper - Claude Code or Cursor via the Apify CLI (latest version,
apify upgrade):apify mcp install claude-code --tools artificially/arxiv-scraper(usecursorinstead ofclaude-codefor Cursor) - Any MCP client (Cursor, VS Code, Windsurf):
{"mcpServers": {"apify": { "url": "https://mcp.apify.com?tools=artificially/arxiv-scraper" }}}
Try asking:
- "Find the 20 most recent arXiv papers on retrieval-augmented generation with citation counts."
- "Search arXiv for papers on AI agents in cs.CL from this month."
Features
Core Search
- Keyword Search: All keywords must match (title, abstract, authors, comments); quote exact phrases or use arXiv syntax (
ti:,au:,abs:,AND,OR) for advanced queries - Category Filtering: Filter by arXiv category (cs.AI, physics, math, etc.)
- Sorting Options: Sort by relevance, submission date, or update date
- Complete Metadata: Title, authors, abstract, categories, dates
Citation Analysis (NEW)
- Citation Counts: Total citations from Semantic Scholar
- Influential Citations: Citations that significantly impacted the field
- Citation Velocity: Recent citation momentum
- Citations Per Year: Historical citation distribution
- Highly Influential Flag: Identify breakthrough papers
Author Profiles (NEW)
- h-Index: Author's impact metric
- Total Citations: Lifetime citation count
- Paper Count: Publication volume
- Affiliations: Current institutional affiliations
- Semantic Scholar Links: Direct profile links
Related Content (NEW)
- References: Papers cited by each result
- Related Papers: AI-recommended similar papers
- Venue Information: Publication venue if applicable
- Fields of Study: Semantic Scholar topic classification
Impact Scoring (NEW)
- Calculated Impact Score: Combined metric considering citations, author h-index, and momentum
- Your sort order is kept: Results follow the
sortByyou choose; sort byimpactScorein the dataset view or export to rank by impact
Use Cases
- Build research paper datasets with citation metrics
- Identify high-impact papers in your field
- Find influential authors and their work
- Track citation trends over time
- Literature review with impact analysis
- Research team evaluation
Input
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
searchQuery | string | Yes | - | Keywords (all must match) or an arXiv query such as ti:transformer AND au:vaswani |
category | string | No | - | arXiv category filter |
maxPapers | number | No | 100 | Maximum papers |
sortBy | string | No | submittedDate | Sort order |
includeCitations | boolean | No | true | Fetch citation metrics |
includeAuthorProfiles | boolean | No | true | Fetch author h-index and stats |
includeReferences | boolean | No | false | Fetch paper bibliography |
maxReferences | number | No | 10 | References per paper |
includeRelatedPapers | boolean | No | false | Fetch similar papers |
maxRelatedPapers | number | No | 5 | Related papers per result |
semanticScholarApiKey | string | No | - | Optional free Semantic Scholar API key for more reliable citation/author data |
compactOutput | boolean | No | false | Slim items with only key fields (see below); recommended for AI agents |
proxyConfiguration | object | No | - | Optional proxy settings |
Example Input
{"searchQuery": "large language models","category": "cs.CL","maxPapers": 50,"includeCitations": true,"includeAuthorProfiles": true,"includeRelatedPapers": true,"sortBy": "submittedDate"}
Output
Each paper produces a result with:
{"arxivId": "2401.12345","title": "Advances in Large Language Models: A Survey","authors": ["John Smith", "Jane Doe"],"authorProfiles": [{"name": "John Smith","authorId": "12345678","hIndex": 45,"citationCount": 15000,"paperCount": 120,"affiliations": ["Stanford University"],"url": "https://www.semanticscholar.org/author/12345678"}],"abstract": "This paper surveys recent advances...","categories": ["cs.CL", "cs.AI"],"categoryDescriptions": ["Computation and Language (NLP)", "Artificial Intelligence"],"citations": {"totalCitations": 1250,"influentialCitations": 89,"citationVelocity": 125.5,"citationsPerYear": {"2023": 450,"2024": 800},"isHighlyInfluential": true},"references": [{"title": "Attention Is All You Need","authors": ["Ashish Vaswani"],"citationCount": 75000,"arxivId": "1706.03762"}],"relatedPapers": [{"title": "GPT-4 Technical Report","citationCount": 5000,"url": "https://arxiv.org/abs/2303.08774"}],"impactScore": 85.3,"venue": "NeurIPS 2024","fieldsOfStudy": ["Computer Science", "Linguistics"],"primaryCategory": "cs.CL","publishedDate": "2024-01-18T17:59:57Z","updatedDate": "2024-01-19T10:02:11Z","doi": null,"journalRef": null,"comments": "25 pages, 6 figures","pdfUrl": "https://arxiv.org/pdf/2401.12345v1","arxivUrl": "https://arxiv.org/abs/2401.12345","semanticScholarId": "204e3073870fae3d05bcbc2f6a8e263d9b72e776","enrichment": "ok","scrapedAt": "2024-01-20T12:00:00Z"}
enrichment tells you what happened with the Semantic Scholar lookup:
ok: citation/author data attachednot_indexed: Semantic Scholar does not know the paper yet (common for papers submitted in the last few days)failed: Semantic Scholar was unavailable (rate limited) after several retries; the paper is still saved with full arXiv metadatadisabled: all enrichment options were turned off
citations, authorProfiles, references and relatedPapers are only present when Semantic Scholar returned data. citationsPerYear and citationVelocity are computed from up to 9,999 of the most recent citing papers that Semantic Scholar returns, so for very highly cited papers they are a sample.
A run summary (SUMMARY) plus lists of failed pages, failed enrichments and skipped entries (FAILED_PAGES, FAILED_ENRICHMENTS, INVALID_PAPERS) are saved to the run's key-value store.
Compact output
With "compactOutput": true each item keeps only the key decision fields: arxivId, title, authors (first 10), authorsCount, primaryCategory, categories, publishedDate, updatedDate, abstractPreview (first 300 characters), doi, journalRef, venue, totalCitations, influentialCitations, impactScore, referencesCount, relatedPapersCount, enrichment, pdfUrl, arxivUrl, scrapedAt. The full abstract, author profiles, reference lists and related paper lists are left out (counts are kept).
Cost
This actor uses pay-per-event pricing; see the Pricing tab for current prices.
You only pay for valid results: entries missing an ID, title, authors or date, duplicates, and failed pages are skipped and never charged. If you set a maximum cost per run, the actor stops cleanly when it is reached.
No API key required - Uses arXiv and Semantic Scholar public APIs. An optional Semantic Scholar API key makes citation data more reliable.
Tips
-
Impact sorting: Sort the dataset by
impactScoreto rank papers by influence (usesortBy: relevanceto get established papers; brand-new submissions are usually not indexed by Semantic Scholar yet) -
Highly influential papers: Look for
isHighlyInfluential: truefor breakthrough papers -
Author quality: Check author h-index to identify papers from established researchers
-
Citation velocity: High velocity indicates trending/hot papers
-
Related papers: Enable
includeRelatedPapersfor comprehensive literature discovery
Rate Limits
- arXiv: 3-second delay between requests, with automatic retries and backoff
- Semantic Scholar: metrics are fetched in batches of 20 papers, with automatic retries and backoff on rate limits
Support
- Built by: Artificially
- Issues: Report bugs or request features via Apify Console
Related actors
- GitHub Repository Scraper - stars, activity, contributors and releases for the code repositories behind papers