arXiv Paper Scraper with Citation Metrics
Pricing
from $2.00 / 1,000 paper records
arXiv Paper Scraper with Citation Metrics
Search arXiv research papers and enrich them with OpenAlex citation counts, FWCI, citing papers and references. Export to CSV, Excel, JSON or XML.
Pricing
from $2.00 / 1,000 paper records
Rating
0.0
(0)
Developer
PunkRecordsData
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
13 hours ago
Last modified
Categories
Share
๐ arXiv Research Papers Scraper - PunkRecordsData
๐ Export arXiv papers with citation metrics in seconds. Search any topic across 2M+ arXiv preprints and get structured rows with title, abstract, authors, categories, PDF link and DOI, enriched with OpenAlex citation counts, FWCI and open-access status, plus optional citing-paper and reference rows. Export to CSV, Excel, JSON or XML.
The arXiv Research Papers Scraper searches arXiv's official API (a single "large language models" query matches 77,764 papers) and joins each result with OpenAlex, the open scholarly graph, so every paper can carry real impact data instead of bare metadata. No API keys needed, and no other arXiv actor on the Apify Store measured in September 2026 offers citation enrichment at all.
| ๐ฏ Target Audience | ๐ก Primary Use Cases |
|---|---|
| AI/ML research teams | Track new papers on a topic with real citation impact |
| R&D and innovation scouts | Rank a field's literature by citations, not just recency |
| Academic librarians and analysts | Build reading lists and bibliometric datasets |
| RAG / LLM pipeline builders | Feed fresh, structured paper metadata into your index |
๐ What the arXiv Paper Scraper does
- arXiv search by keyword over all fields, title, abstract, author or comment, with the full official 155-category taxonomy as a filter and relevance/date sorting.
- Direct lookups by arXiv ID (e.g. 1706.03762) in the same run.
- Citation metrics per paper (OpenAlex): citation count, FWCI (field-weighted citation impact), open-access status, top concepts and the OpenAlex ID. Charged only when a confident title match is found.
- Citing papers: one row per paper that cites your result, newest first, with DOI, venue, authors and its own citation count.
- References: one row per work a paper cites, resolved to full OpenAlex records.
๐ก Why it matters: raw arXiv metadata tells you a paper exists; the citation layer tells you whether it matters. Getting both in one dataset normally means writing your own two-API join. Here it is one input form.
๐ Output of the arXiv paper search
One row per paper (20 fields), plus optional citation and reference rows. Real sample from a live run:
{"recordType": "paper","arxivId": "1706.03762v7","title": "Attention Is All You Need","url": "https://arxiv.org/abs/1706.03762v7","pdfUrl": "https://arxiv.org/pdf/1706.03762v7","authors": ["Ashish Vaswani", "Noam Shazeer", "Niki Parmar", "Jakob Uszkoreit"],"primaryCategory": "cs.CL","published": "2017-06-12T17:57:34Z","citationCount": 26892,"openAccessStatus": "green","topConcepts": ["Computer science", "Artificial intelligence", "Machine translation"],"openAlexId": "W2626778328","error": null}
{"recordType": "reference","sourceArxivId": "1706.03762v7","title": "Effective Approaches to Attention-based Neural Machine Translation","openAlexId": "W1902237438","doi": "10.18653/v1/d15-1166","citedByCount": 8655,"error": null}
โจ Why choose this arXiv scraper
- The only measured arXiv actor with citation enrichment. Others return search metadata; this one adds impact numbers per paper.
- 4 billable events, each switchable: papers, metrics, citing papers, references. Pay only for the depth you use.
- Full official category taxonomy (155 categories) as a real dropdown filter, not free text you have to guess.
- Polite by design: respects arXiv's 3-second request guidance and identifies itself to OpenAlex, so runs are stable at volume.
- Honest billing: metrics are charged only when the OpenAlex match is confident; unmatched papers show "Not Available" and cost nothing extra.
๐ How this arXiv scraper compares to alternatives
Measured against the arXiv actors on the Apify Store (September 2026):
| This actor | Other arXiv actors | |
|---|---|---|
| Citation count / FWCI per paper | Yes (OpenAlex) | No |
| Citing papers and references as rows | Yes, separate events | No |
| Billable data events | 4 | 1 |
| Category filter | Full 155-category dropdown | Free text or none |
| Price per 1,000 papers | $2.50 | $2.00 (closest priced alternative) |
๐ How to use the arXiv Research Papers Scraper
- Create a free Apify account (it includes $5 of platform credit) at console.apify.com.
- Open this actor's page and click Try for free.
- Enter search terms (or paste specific arXiv IDs).
- Optionally narrow by category, switch on citing papers or references, and set Max Items.
- Click Start and download the dataset as CSV, Excel, JSON or XML.
๐ผ Business use cases
Competitive R&D monitoring
Weekly scheduled run on your field's keywords, sorted by submission date, with citation counts to separate noise from signal.
Literature reviews at scale
Pull hundreds of papers with abstracts, categories and impact metrics into one spreadsheet instead of copy-pasting from the site.
Building citation graphs
Enable citing papers and references to export edge lists (paper โ cited-by, paper โ references) ready for network analysis.
Powering AI research assistants
Feed structured, current paper metadata into RAG pipelines; the PDF links come resolved per version.
๐ Automating the arXiv Paper Scraper
Connect to Make, Zapier, Slack, Airbyte, GitHub or Google Drive via Apify's integrations: schedule weekly topic sweeps, push new high-citation papers to a Slack channel, or sync results into your data warehouse.
๐ Beyond business use cases
- Research: bibliometric studies with FWCI and concept tagging included.
- Personal: track when your own papers get new citations.
- Non-profit: open-science monitoring with open-access status per paper.
- Experimentation: a zero-setup playground for the arXiv + OpenAlex APIs.
๐ค Ask an AI assistant about this scraper
"I need arXiv papers on a topic as structured data with citation counts, plus the papers that cite them. Would the arXiv Research Papers Scraper on Apify (apify.com/punkrecordsdata/arxiv-research-papers-scraper) cover this, and how would I schedule it weekly?"
โ Frequently Asked Questions
๐ How do I search arXiv papers by keyword and export to CSV?
Enter your terms, pick the field to search (title, abstract, author, all), click Start, and download the dataset as CSV from the Storage tab.
๐ How do I get citation counts for arXiv papers?
Leave the "Citation metrics" module on. Each paper is matched against OpenAlex and enriched with citation count, FWCI and open-access status when the match is confident.
๐ธ Can I get the papers that cite a specific paper?
Yes. Paste the arXiv ID, enable "Citing papers", and each citing work becomes its own dataset row with DOI, venue and date.
๐ Can I fetch specific papers by arXiv ID?
Yes, the "Specific arXiv IDs" field accepts a list (with or without version suffix) and runs alongside any search terms.
๐ท Which categories can I filter by?
All 155 official arXiv categories, from cs.AI to q-fin.TR, as a multi-select dropdown.
๐ต Do I pay for metrics when no match is found?
No. The metrics event is charged only when OpenAlex returns a confident match; otherwise the fields read "Not Available" at no extra cost.
๐ Does it download the PDFs?
It returns the direct PDF URL per paper version; downloading the files themselves is up to your pipeline.
โฑ How fast is it?
arXiv asks clients for one request every 3 seconds, which the actor respects; a 100-paper page costs one request, so 1,000 papers take about half a minute of API time plus enrichment.
๐งฎ Why do citation numbers differ from Google Scholar?
Counts come from OpenAlex, which is open and reproducible; Google Scholar typically shows higher, non-reproducible counts. The OpenAlex ID is included so you can verify every number.
๐ฆ Can I analyze thousands of papers in one run?
Yes, paid plans allow up to 1,000,000 papers per run. Free users get a 10-paper preview.
โ๏ธ What happens if OpenAlex is briefly rate-limited?
The actor retries politely; if a paper still cannot be enriched, it ships with metadata only and you are not charged for metrics on it.
๐ Integrate with any app
Every dataset is available through the Apify API in JSON, CSV, Excel or XML for Python, R, Node.js, Google Sheets or any HTTP client. Webhooks fire when a run finishes.
๐ Recommended Actors
- SEC EDGAR Filings Scraper - official filings with the same structured rigor
- USPTO Trademark Status Scraper - IP records for innovation research
- Grants.gov Opportunities Scraper - federal funding to pair with the literature
- OpenCV Image Analyzer - computer vision on figures and images
๐ก Pro Tip: browse the complete PunkRecordsData collection for more data tools.
๐ Need Help? contact.punkrecordsdata@gmail.com
โ ๏ธ Disclaimer: independent tool, not affiliated with arXiv or OpenAlex; only publicly available data. Thank you to arXiv for use of its open access interoperability.