arXiv Scraper avatar

arXiv Scraper

Deprecated

Pricing

from $0.70 / 1,000 results

Go to Apify Store
arXiv Scraper

arXiv Scraper

Deprecated

Search and extract academic papers from arXiv.org. Get paper titles, authors, abstracts, categories, and PDF links for AI/ML, physics, math, and more.

Pricing

from $0.70 / 1,000 results

Rating

0.0

(0)

Developer

Artificially

Artificially

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

5 days ago

Last modified

Share

arXiv Papers Scraper - Enhanced

Search and extract academic papers from arXiv.org with citation analysis, author profiles, and impact metrics via Semantic Scholar integration.

Use with AI agents (MCP)

This actor works as a tool for AI agents through Apify's MCP server at https://mcp.apify.com, so Claude, ChatGPT, Cursor and other MCP clients can search arXiv papers directly. For agents, set "compactOutput": true and keep maxPapers small (5-20) to get short, token-friendly results.

Quick setup (sign in with your Apify account when asked):

  • Claude (claude.ai or Claude Desktop): Settings → Connectors → Add custom connector, and paste https://mcp.apify.com?tools=artificially/arxiv-scraper
  • Claude Code or Cursor via the Apify CLI (latest version, apify upgrade): apify mcp install claude-code --tools artificially/arxiv-scraper (use cursor instead of claude-code for Cursor)
  • Any MCP client (Cursor, VS Code, Windsurf):
{
"mcpServers": {
"apify": { "url": "https://mcp.apify.com?tools=artificially/arxiv-scraper" }
}
}

Try asking:

  • "Find the 20 most recent arXiv papers on retrieval-augmented generation with citation counts."
  • "Search arXiv for papers on AI agents in cs.CL from this month."

Features

  • Keyword Search: All keywords must match (title, abstract, authors, comments); quote exact phrases or use arXiv syntax (ti:, au:, abs:, AND, OR) for advanced queries
  • Category Filtering: Filter by arXiv category (cs.AI, physics, math, etc.)
  • Sorting Options: Sort by relevance, submission date, or update date
  • Complete Metadata: Title, authors, abstract, categories, dates

Citation Analysis (NEW)

  • Citation Counts: Total citations from Semantic Scholar
  • Influential Citations: Citations that significantly impacted the field
  • Citation Velocity: Recent citation momentum
  • Citations Per Year: Historical citation distribution
  • Highly Influential Flag: Identify breakthrough papers

Author Profiles (NEW)

  • h-Index: Author's impact metric
  • Total Citations: Lifetime citation count
  • Paper Count: Publication volume
  • Affiliations: Current institutional affiliations
  • Semantic Scholar Links: Direct profile links
  • References: Papers cited by each result
  • Related Papers: AI-recommended similar papers
  • Venue Information: Publication venue if applicable
  • Fields of Study: Semantic Scholar topic classification

Impact Scoring (NEW)

  • Calculated Impact Score: Combined metric considering citations, author h-index, and momentum
  • Your sort order is kept: Results follow the sortBy you choose; sort by impactScore in the dataset view or export to rank by impact

Use Cases

  • Build research paper datasets with citation metrics
  • Identify high-impact papers in your field
  • Find influential authors and their work
  • Track citation trends over time
  • Literature review with impact analysis
  • Research team evaluation

Input

FieldTypeRequiredDefaultDescription
searchQuerystringYes-Keywords (all must match) or an arXiv query such as ti:transformer AND au:vaswani
categorystringNo-arXiv category filter
maxPapersnumberNo100Maximum papers
sortBystringNosubmittedDateSort order
includeCitationsbooleanNotrueFetch citation metrics
includeAuthorProfilesbooleanNotrueFetch author h-index and stats
includeReferencesbooleanNofalseFetch paper bibliography
maxReferencesnumberNo10References per paper
includeRelatedPapersbooleanNofalseFetch similar papers
maxRelatedPapersnumberNo5Related papers per result
semanticScholarApiKeystringNo-Optional free Semantic Scholar API key for more reliable citation/author data
compactOutputbooleanNofalseSlim items with only key fields (see below); recommended for AI agents
proxyConfigurationobjectNo-Optional proxy settings

Example Input

{
"searchQuery": "large language models",
"category": "cs.CL",
"maxPapers": 50,
"includeCitations": true,
"includeAuthorProfiles": true,
"includeRelatedPapers": true,
"sortBy": "submittedDate"
}

Output

Each paper produces a result with:

{
"arxivId": "2401.12345",
"title": "Advances in Large Language Models: A Survey",
"authors": ["John Smith", "Jane Doe"],
"authorProfiles": [
{
"name": "John Smith",
"authorId": "12345678",
"hIndex": 45,
"citationCount": 15000,
"paperCount": 120,
"affiliations": ["Stanford University"],
"url": "https://www.semanticscholar.org/author/12345678"
}
],
"abstract": "This paper surveys recent advances...",
"categories": ["cs.CL", "cs.AI"],
"categoryDescriptions": ["Computation and Language (NLP)", "Artificial Intelligence"],
"citations": {
"totalCitations": 1250,
"influentialCitations": 89,
"citationVelocity": 125.5,
"citationsPerYear": {
"2023": 450,
"2024": 800
},
"isHighlyInfluential": true
},
"references": [
{
"title": "Attention Is All You Need",
"authors": ["Ashish Vaswani"],
"citationCount": 75000,
"arxivId": "1706.03762"
}
],
"relatedPapers": [
{
"title": "GPT-4 Technical Report",
"citationCount": 5000,
"url": "https://arxiv.org/abs/2303.08774"
}
],
"impactScore": 85.3,
"venue": "NeurIPS 2024",
"fieldsOfStudy": ["Computer Science", "Linguistics"],
"primaryCategory": "cs.CL",
"publishedDate": "2024-01-18T17:59:57Z",
"updatedDate": "2024-01-19T10:02:11Z",
"doi": null,
"journalRef": null,
"comments": "25 pages, 6 figures",
"pdfUrl": "https://arxiv.org/pdf/2401.12345v1",
"arxivUrl": "https://arxiv.org/abs/2401.12345",
"semanticScholarId": "204e3073870fae3d05bcbc2f6a8e263d9b72e776",
"enrichment": "ok",
"scrapedAt": "2024-01-20T12:00:00Z"
}

enrichment tells you what happened with the Semantic Scholar lookup:

  • ok: citation/author data attached
  • not_indexed: Semantic Scholar does not know the paper yet (common for papers submitted in the last few days)
  • failed: Semantic Scholar was unavailable (rate limited) after several retries; the paper is still saved with full arXiv metadata
  • disabled: all enrichment options were turned off

citations, authorProfiles, references and relatedPapers are only present when Semantic Scholar returned data. citationsPerYear and citationVelocity are computed from up to 9,999 of the most recent citing papers that Semantic Scholar returns, so for very highly cited papers they are a sample.

A run summary (SUMMARY) plus lists of failed pages, failed enrichments and skipped entries (FAILED_PAGES, FAILED_ENRICHMENTS, INVALID_PAPERS) are saved to the run's key-value store.

Compact output

With "compactOutput": true each item keeps only the key decision fields: arxivId, title, authors (first 10), authorsCount, primaryCategory, categories, publishedDate, updatedDate, abstractPreview (first 300 characters), doi, journalRef, venue, totalCitations, influentialCitations, impactScore, referencesCount, relatedPapersCount, enrichment, pdfUrl, arxivUrl, scrapedAt. The full abstract, author profiles, reference lists and related paper lists are left out (counts are kept).

Cost

This actor uses pay-per-event pricing; see the Pricing tab for current prices.

You only pay for valid results: entries missing an ID, title, authors or date, duplicates, and failed pages are skipped and never charged. If you set a maximum cost per run, the actor stops cleanly when it is reached.

No API key required - Uses arXiv and Semantic Scholar public APIs. An optional Semantic Scholar API key makes citation data more reliable.

Tips

  1. Impact sorting: Sort the dataset by impactScore to rank papers by influence (use sortBy: relevance to get established papers; brand-new submissions are usually not indexed by Semantic Scholar yet)

  2. Highly influential papers: Look for isHighlyInfluential: true for breakthrough papers

  3. Author quality: Check author h-index to identify papers from established researchers

  4. Citation velocity: High velocity indicates trending/hot papers

  5. Related papers: Enable includeRelatedPapers for comprehensive literature discovery

Rate Limits

  • arXiv: 3-second delay between requests, with automatic retries and backoff
  • Semantic Scholar: metrics are fetched in batches of 20 papers, with automatic retries and backoff on rate limits

Support

  • Built by: Artificially
  • Issues: Report bugs or request features via Apify Console