arXiv Paper Scraper with Citation Metrics avatar

arXiv Paper Scraper with Citation Metrics

Pricing

from $12.00 / 1,000 paper records

Go to Apify Store
arXiv Paper Scraper with Citation Metrics

arXiv Paper Scraper with Citation Metrics

Search arXiv research papers by keyword, author, category or arXiv ID and export title, abstract, authors, PDF link and DOI, plus OpenAlex citation counts, FWCI and references. Export CSV, Excel, JSON. No login or API key. Filters all 155 arXiv categories.

Pricing

from $12.00 / 1,000 paper records

Rating

0.0

(0)

Developer

RecordsData

RecordsData

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

18 hours ago

Last modified

Categories

Share

PunkRecordsData

๐Ÿ“š arXiv Scraper: Papers, Citations & Impact Metrics

arXiv Research Papers Scraper searches arXiv by keyword, author, category or arXiv ID and returns one row per paper: title, abstract, authors, categories, PDF link, DOI and dates. Each paper can be enriched with OpenAlex citation counts, FWCI and open-access status, plus citing-paper and reference rows. Export CSV, Excel, JSON or XML. No login, no API key. Verified against live arXiv on 2026-10-04: 10 papers in 14 seconds.

The arXiv Research Papers Scraper queries arXiv's official API (the query "large language models" matches 78,607 papers on 2026-10-04) and joins every result with OpenAlex, the open scholarly graph, so a paper row can carry impact data instead of bare metadata. It is built for research teams, analysts, librarians and RAG pipeline builders who need fresh, structured paper data without writing a two-API join.

๐Ÿ”Ž What does the arXiv Scraper do?

  • Search arXiv papers by keyword across all fields, title, abstract, author, category or comment.
  • Fetch papers by arXiv ID (for example 1706.03762) in the same run.
  • Filter by category with a dropdown of all 155 official arXiv categories (cs.AI, cs.CL, stat.ML, quant-ph and more).
  • Sort by relevance, last updated date or submitted date, ascending or descending.
  • Add citation metrics from OpenAlex: citation count, FWCI, open-access status, top concepts.
  • Export citing papers (who cites a paper) and references (what a paper cites) as separate rows.

๐Ÿ“Š What data can you extract from arXiv?

FieldDescription
recordTypepaper, citation or reference
arxivId, url, pdfUrlarXiv ID with version, abstract page and direct PDF link
title, summaryPaper title and full abstract
authorsAuthor list
primaryCategory, categoriesMain arXiv category and all categories
published, updatedFirst submission and last update timestamps
doi, journalRef, commentDOI, journal reference and author comment (Not Disclosed when arXiv has none)
citationCount, fwci, openAccessStatus, openAlexIdOpenAlex impact data (Not Available when no confident match)
topConceptsUp to 5 OpenAlex concepts, present only when OpenAlex returns them
searchTermThe search term that found the paper (absent for ID lookups)
sourceArxivId, sourceTitle, publicationDate, citedByCount, venueOn citation and reference rows
scrapedAtRun timestamp (UTC)

๐Ÿงช Sample output from a real run

Input: arxivIds: ["1706.03762"] with citing papers and references on. Abstract shortened here; the dataset has the full text.

{
"recordType": "paper",
"arxivId": "1706.03762v7",
"title": "Attention Is All You Need",
"url": "https://arxiv.org/abs/1706.03762v7",
"pdfUrl": "https://arxiv.org/pdf/1706.03762v7",
"summary": "The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. ...",
"authors": ["Ashish Vaswani", "Noam Shazeer", "Niki Parmar", "Jakob Uszkoreit", "Llion Jones", "Aidan N. Gomez", "Lukasz Kaiser", "Illia Polosukhin"],
"primaryCategory": "cs.CL",
"categories": ["cs.CL", "cs.LG"],
"published": "2017-06-12T17:57:34Z",
"updated": "2023-08-02T00:41:18Z",
"doi": "Not Disclosed",
"journalRef": "Not Disclosed",
"comment": "15 pages, 5 figures",
"citationCount": 26739,
"fwci": "Not Disclosed",
"openAccessStatus": "gold",
"openAlexId": "W2626778328",
"topConcepts": ["Computer science", "Machine translation", "Transformer", "BLEU", "Encoder"],
"scrapedAt": "2026-10-04T05:03:31.911Z"
}

A reference row from the same run:

{
"recordType": "reference",
"sourceArxivId": "1706.03762v7",
"sourceTitle": "Attention Is All You Need",
"title": "Building a Large Annotated Corpus of English: The Penn Treebank",
"openAlexId": "W1632114991",
"doi": "10.21236/ada273556",
"publicationDate": "1993-04-30",
"citedByCount": 7485,
"authors": ["Mitchell P. Marcus"],
"venue": "Not Disclosed",
"scrapedAt": "2026-10-04T05:03:34.579Z"
}

๐Ÿ’ต How much does it cost to scrape arXiv papers?

Pay per event, no subscription. Prices at the Free tier (volume tiers are cheaper on the actor page):

EventPriceWhen it is charged
Paper record$0.015One arXiv paper saved to the dataset
Citation metrics (OpenAlex)$0.012Only when OpenAlex returns a confident title match
Citing paper$0.010One citing-paper row (off by default)
Reference$0.010One reference row (off by default)

Measured example: a 10-paper run for "large language models" delivered 10 papers, 9 of them matched in OpenAlex, so it charged 10 paper events and 9 metrics events (about $0.26 at the prices above). With metrics switched off, 1,000 papers cost about $15. You are never charged for rows that failed, for unknown arXiv IDs, for empty searches or for papers OpenAlex could not match. Set a maximum charge per run and the actor stops cleanly when it is reached. Free Apify users get a 10-paper preview.

โ–ถ๏ธ How to scrape arXiv papers in 3 steps

  1. Open the actor and click Try for free.
  2. Enter search terms (or arXiv IDs), optionally pick categories, and set Max Items.
  3. Click Start, then download the dataset as CSV, Excel, JSON or XML.

โš™๏ธ Input

FieldDefaultWhat it does
searchTerms["large language models"]Keywords, one search per term
arxivIds[]Exact arXiv IDs to fetch directly
maxItems10Maximum papers (free users capped at 10, paid up to 1,000,000)
searchFieldallall, ti, abs, au, cat or co
categories[]One or more arXiv categories, combined with AND
sortBy / sortOrderrelevance / descendingAlso lastUpdatedDate, submittedDate
includeMetricstrueOpenAlex citation count, FWCI, open access, concepts
includeCitations, maxCitationsPerPaperfalse, 25Citing-paper rows per paper (max 200)
includeReferences, maxReferencesPerPaperfalse, 25Reference rows per paper (max 200)
{
"searchTerms": ["quantum error correction"],
"categories": ["quant-ph"],
"sortBy": "submittedDate",
"sortOrder": "descending",
"maxItems": 50,
"includeMetrics": true
}

๐Ÿ“ค Output

The dataset has one row per paper, citing paper or reference, told apart by recordType. The Overview view shows the key columns; the full view has every field. Optional fields (such as topConcepts) are left out when empty instead of being shipped as blank columns. If a search finds nothing, the run ends as succeeded with the message "No arXiv papers matched the input. Nothing was charged."

โš–๏ธ arXiv scraper vs alternatives

Prices read from the public Apify Store API for 16 arXiv-related actors on 2026-10-04:

This actorTypical metadata-only arXiv actor
Price per paper$0.015 (+ $0.012 metrics)$0.00125 to $0.003 per result, plus a small start fee
Citation count and FWCI per paperYes, via OpenAlexNot offered by the actors checked that return only paper metadata
Citing papers and references as rowsYes, separate eventsNot offered by those actors
Category filterDropdown of 155 categoriesVaries

We are more expensive per paper. If you only need titles and abstracts in bulk, a metadata-only actor is cheaper. Choose this one when citation counts, citing papers or references matter, or when you want one dataset instead of joining arXiv and OpenAlex yourself.

๐Ÿ’ผ Use cases

  • R&D monitoring: schedule a weekly run on your field's keywords, sorted by submission date, with citation counts to separate signal from noise.
  • Literature reviews: pull hundreds of papers with abstracts, categories and impact metrics into one spreadsheet.
  • Citation graphs: turn on citing papers and references to export edge lists for network analysis.
  • AI research assistants and RAG: feed structured, current paper metadata and PDF links into your index.

๐Ÿ”Œ Run the arXiv scraper via API and integrations

curl -X POST "https://api.apify.com/v2/acts/RecordsData~arxiv-research-papers-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"searchTerms":["retrieval augmented generation"],"maxItems":20}'

Use the Apify API, the Python or JavaScript client, schedules and webhooks, or connect Zapier, Make, n8n, Google Sheets, Slack and Airbyte. The actor is also available to AI agents through the Apify MCP server.

The actor reads public paper metadata through arXiv's official API and OpenAlex's open API. It waits about 3 seconds between arXiv requests, as arXiv asks, and identifies itself with a contact address. It does not download PDFs, only links to them. Check arXiv's terms and your own jurisdiction's rules for how you reuse abstracts. Independent tool, not affiliated with arXiv or OpenAlex. Thank you to arXiv for use of its open access interoperability.

โ“ Frequently asked questions

How do I search arXiv papers by keyword and export to CSV?

Enter your terms, pick the field to search, click Start and download the dataset as CSV from the Storage tab.

How do I get citation counts for arXiv papers?

Leave Citation metrics on. Each paper is matched by title against OpenAlex and gets citation count, FWCI, open-access status and concepts when the match is confident.

Why are some fields "Not Disclosed" or "Not Available"?

Not Disclosed means the source has no value (many arXiv papers have no DOI or journal reference, and OpenAlex has no FWCI for very recent papers). Not Available means OpenAlex had no confident match for that paper. In a 10-paper run, 9 matched.

Can I get the papers that cite a specific paper?

Yes. Add the arXiv ID, turn on Citing papers, and each citing work becomes its own row, newest first.

What happens if I enter a wrong arXiv ID?

arXiv rejects it, the actor logs a warning and skips it, and nothing is charged. Valid IDs in the same list still resolve.

Why did I get 0 results?

The search matched nothing on arXiv (try fewer words, or search "all" fields), or the category filter is too narrow. The run ends as succeeded with zero charge. If arXiv itself is unreachable and no paper was delivered, the run fails with a clear message and charges nothing.

Does it download the PDFs?

No. It returns the direct PDF URL for each paper version.

How fast is it?

A 10-paper run with metrics finished in 14 seconds. arXiv asks for one request every 3 seconds, so large runs are paced by that.

Why do citation numbers differ from Google Scholar?

Counts come from OpenAlex, which is open and reproducible. The OpenAlex ID is in every row so you can check any number.

Can I analyze thousands of papers in one run?

Yes, paid plans allow up to 1,000,000 papers per run. Use maxItems and the maximum charge setting to control spend.

๐Ÿ†˜ Support

Questions or a field you need? Open the Issues tab on the actor page or write to contact.punkrecordsdata@gmail.com.

Last updated: 2026-10-03