arXiv Paper Scraper with Citation Metrics
Pricing
from $12.00 / 1,000 paper records
arXiv Paper Scraper with Citation Metrics
Search arXiv research papers by keyword, author, category or arXiv ID and export title, abstract, authors, PDF link and DOI, plus OpenAlex citation counts, FWCI and references. Export CSV, Excel, JSON. No login or API key. Filters all 155 arXiv categories.
Pricing
from $12.00 / 1,000 paper records
Rating
0.0
(0)
Developer
RecordsData
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
18 hours ago
Last modified
Categories
Share
๐ arXiv Scraper: Papers, Citations & Impact Metrics
arXiv Research Papers Scraper searches arXiv by keyword, author, category or arXiv ID and returns one row per paper: title, abstract, authors, categories, PDF link, DOI and dates. Each paper can be enriched with OpenAlex citation counts, FWCI and open-access status, plus citing-paper and reference rows. Export CSV, Excel, JSON or XML. No login, no API key. Verified against live arXiv on 2026-10-04: 10 papers in 14 seconds.
The arXiv Research Papers Scraper queries arXiv's official API (the query "large language models" matches 78,607 papers on 2026-10-04) and joins every result with OpenAlex, the open scholarly graph, so a paper row can carry impact data instead of bare metadata. It is built for research teams, analysts, librarians and RAG pipeline builders who need fresh, structured paper data without writing a two-API join.
๐ What does the arXiv Scraper do?
- Search arXiv papers by keyword across all fields, title, abstract, author, category or comment.
- Fetch papers by arXiv ID (for example
1706.03762) in the same run. - Filter by category with a dropdown of all 155 official arXiv categories (cs.AI, cs.CL, stat.ML, quant-ph and more).
- Sort by relevance, last updated date or submitted date, ascending or descending.
- Add citation metrics from OpenAlex: citation count, FWCI, open-access status, top concepts.
- Export citing papers (who cites a paper) and references (what a paper cites) as separate rows.
๐ What data can you extract from arXiv?
| Field | Description |
|---|---|
recordType | paper, citation or reference |
arxivId, url, pdfUrl | arXiv ID with version, abstract page and direct PDF link |
title, summary | Paper title and full abstract |
authors | Author list |
primaryCategory, categories | Main arXiv category and all categories |
published, updated | First submission and last update timestamps |
doi, journalRef, comment | DOI, journal reference and author comment (Not Disclosed when arXiv has none) |
citationCount, fwci, openAccessStatus, openAlexId | OpenAlex impact data (Not Available when no confident match) |
topConcepts | Up to 5 OpenAlex concepts, present only when OpenAlex returns them |
searchTerm | The search term that found the paper (absent for ID lookups) |
sourceArxivId, sourceTitle, publicationDate, citedByCount, venue | On citation and reference rows |
scrapedAt | Run timestamp (UTC) |
๐งช Sample output from a real run
Input: arxivIds: ["1706.03762"] with citing papers and references on. Abstract shortened here; the dataset has the full text.
{"recordType": "paper","arxivId": "1706.03762v7","title": "Attention Is All You Need","url": "https://arxiv.org/abs/1706.03762v7","pdfUrl": "https://arxiv.org/pdf/1706.03762v7","summary": "The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. ...","authors": ["Ashish Vaswani", "Noam Shazeer", "Niki Parmar", "Jakob Uszkoreit", "Llion Jones", "Aidan N. Gomez", "Lukasz Kaiser", "Illia Polosukhin"],"primaryCategory": "cs.CL","categories": ["cs.CL", "cs.LG"],"published": "2017-06-12T17:57:34Z","updated": "2023-08-02T00:41:18Z","doi": "Not Disclosed","journalRef": "Not Disclosed","comment": "15 pages, 5 figures","citationCount": 26739,"fwci": "Not Disclosed","openAccessStatus": "gold","openAlexId": "W2626778328","topConcepts": ["Computer science", "Machine translation", "Transformer", "BLEU", "Encoder"],"scrapedAt": "2026-10-04T05:03:31.911Z"}
A reference row from the same run:
{"recordType": "reference","sourceArxivId": "1706.03762v7","sourceTitle": "Attention Is All You Need","title": "Building a Large Annotated Corpus of English: The Penn Treebank","openAlexId": "W1632114991","doi": "10.21236/ada273556","publicationDate": "1993-04-30","citedByCount": 7485,"authors": ["Mitchell P. Marcus"],"venue": "Not Disclosed","scrapedAt": "2026-10-04T05:03:34.579Z"}
๐ต How much does it cost to scrape arXiv papers?
Pay per event, no subscription. Prices at the Free tier (volume tiers are cheaper on the actor page):
| Event | Price | When it is charged |
|---|---|---|
| Paper record | $0.015 | One arXiv paper saved to the dataset |
| Citation metrics (OpenAlex) | $0.012 | Only when OpenAlex returns a confident title match |
| Citing paper | $0.010 | One citing-paper row (off by default) |
| Reference | $0.010 | One reference row (off by default) |
Measured example: a 10-paper run for "large language models" delivered 10 papers, 9 of them matched in OpenAlex, so it charged 10 paper events and 9 metrics events (about $0.26 at the prices above). With metrics switched off, 1,000 papers cost about $15. You are never charged for rows that failed, for unknown arXiv IDs, for empty searches or for papers OpenAlex could not match. Set a maximum charge per run and the actor stops cleanly when it is reached. Free Apify users get a 10-paper preview.
โถ๏ธ How to scrape arXiv papers in 3 steps
- Open the actor and click Try for free.
- Enter search terms (or arXiv IDs), optionally pick categories, and set Max Items.
- Click Start, then download the dataset as CSV, Excel, JSON or XML.
โ๏ธ Input
| Field | Default | What it does |
|---|---|---|
searchTerms | ["large language models"] | Keywords, one search per term |
arxivIds | [] | Exact arXiv IDs to fetch directly |
maxItems | 10 | Maximum papers (free users capped at 10, paid up to 1,000,000) |
searchField | all | all, ti, abs, au, cat or co |
categories | [] | One or more arXiv categories, combined with AND |
sortBy / sortOrder | relevance / descending | Also lastUpdatedDate, submittedDate |
includeMetrics | true | OpenAlex citation count, FWCI, open access, concepts |
includeCitations, maxCitationsPerPaper | false, 25 | Citing-paper rows per paper (max 200) |
includeReferences, maxReferencesPerPaper | false, 25 | Reference rows per paper (max 200) |
{"searchTerms": ["quantum error correction"],"categories": ["quant-ph"],"sortBy": "submittedDate","sortOrder": "descending","maxItems": 50,"includeMetrics": true}
๐ค Output
The dataset has one row per paper, citing paper or reference, told apart by recordType. The Overview view shows the key columns; the full view has every field. Optional fields (such as topConcepts) are left out when empty instead of being shipped as blank columns. If a search finds nothing, the run ends as succeeded with the message "No arXiv papers matched the input. Nothing was charged."
โ๏ธ arXiv scraper vs alternatives
Prices read from the public Apify Store API for 16 arXiv-related actors on 2026-10-04:
| This actor | Typical metadata-only arXiv actor | |
|---|---|---|
| Price per paper | $0.015 (+ $0.012 metrics) | $0.00125 to $0.003 per result, plus a small start fee |
| Citation count and FWCI per paper | Yes, via OpenAlex | Not offered by the actors checked that return only paper metadata |
| Citing papers and references as rows | Yes, separate events | Not offered by those actors |
| Category filter | Dropdown of 155 categories | Varies |
We are more expensive per paper. If you only need titles and abstracts in bulk, a metadata-only actor is cheaper. Choose this one when citation counts, citing papers or references matter, or when you want one dataset instead of joining arXiv and OpenAlex yourself.
๐ผ Use cases
- R&D monitoring: schedule a weekly run on your field's keywords, sorted by submission date, with citation counts to separate signal from noise.
- Literature reviews: pull hundreds of papers with abstracts, categories and impact metrics into one spreadsheet.
- Citation graphs: turn on citing papers and references to export edge lists for network analysis.
- AI research assistants and RAG: feed structured, current paper metadata and PDF links into your index.
๐ Run the arXiv scraper via API and integrations
curl -X POST "https://api.apify.com/v2/acts/RecordsData~arxiv-research-papers-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"searchTerms":["retrieval augmented generation"],"maxItems":20}'
Use the Apify API, the Python or JavaScript client, schedules and webhooks, or connect Zapier, Make, n8n, Google Sheets, Slack and Airbyte. The actor is also available to AI agents through the Apify MCP server.
๐ก๏ธ Is it legal to scrape arXiv?
The actor reads public paper metadata through arXiv's official API and OpenAlex's open API. It waits about 3 seconds between arXiv requests, as arXiv asks, and identifies itself with a contact address. It does not download PDFs, only links to them. Check arXiv's terms and your own jurisdiction's rules for how you reuse abstracts. Independent tool, not affiliated with arXiv or OpenAlex. Thank you to arXiv for use of its open access interoperability.
โ Frequently asked questions
How do I search arXiv papers by keyword and export to CSV?
Enter your terms, pick the field to search, click Start and download the dataset as CSV from the Storage tab.
How do I get citation counts for arXiv papers?
Leave Citation metrics on. Each paper is matched by title against OpenAlex and gets citation count, FWCI, open-access status and concepts when the match is confident.
Why are some fields "Not Disclosed" or "Not Available"?
Not Disclosed means the source has no value (many arXiv papers have no DOI or journal reference, and OpenAlex has no FWCI for very recent papers). Not Available means OpenAlex had no confident match for that paper. In a 10-paper run, 9 matched.
Can I get the papers that cite a specific paper?
Yes. Add the arXiv ID, turn on Citing papers, and each citing work becomes its own row, newest first.
What happens if I enter a wrong arXiv ID?
arXiv rejects it, the actor logs a warning and skips it, and nothing is charged. Valid IDs in the same list still resolve.
Why did I get 0 results?
The search matched nothing on arXiv (try fewer words, or search "all" fields), or the category filter is too narrow. The run ends as succeeded with zero charge. If arXiv itself is unreachable and no paper was delivered, the run fails with a clear message and charges nothing.
Does it download the PDFs?
No. It returns the direct PDF URL for each paper version.
How fast is it?
A 10-paper run with metrics finished in 14 seconds. arXiv asks for one request every 3 seconds, so large runs are paced by that.
Why do citation numbers differ from Google Scholar?
Counts come from OpenAlex, which is open and reproducible. The OpenAlex ID is in every row so you can check any number.
Can I analyze thousands of papers in one run?
Yes, paid plans allow up to 1,000,000 papers per run. Use maxItems and the maximum charge setting to control spend.
๐ Want related data? Other PunkRecordsData scrapers
- PubMed Scraper for biomedical literature
- SEC EDGAR Filings Scraper for official company filings
- Grants.gov Opportunities Scraper for federal funding to pair with the literature
๐ Support
Questions or a field you need? Open the Issues tab on the actor page or write to contact.punkrecordsdata@gmail.com.
Last updated: 2026-10-03