arXiv Scraper - Papers, Abstracts & PDF Links
Pricing
from $1.94 / 1,000 papers
arXiv Scraper - Papers, Abstracts & PDF Links
Search arXiv papers by title, author, abstract or category, in arXiv's own query syntax. Each row holds the arXiv id, title, full abstract, authors, categories, first-submitted and last-revised dates, the DOI when there is one, and abstract and PDF links. $2.00 per 1,000 papers.
Pricing
from $1.94 / 1,000 papers
Rating
5.0
(1)
Developer
Dami's Studio
Maintained by CommunityActor stats
0
Bookmarked
4
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Write an arXiv query and get one row per paper: the arXiv id, the title, the whole abstract, every author, the categories, the first-submitted and last-revised dates, the DOI when there is one, and direct links to the abstract page and the PDF.
Metadata and links only. The PDF link is there for you to fetch; this actor does not download or read the paper, and arXiv publishes no citation counts, so there are none here either.
| Input | An arXiv search query, using arXiv's own field prefixes |
| Output | One row per paper: id, title, full abstract, authors, categories, both dates, DOI, abstract and PDF links |
| Ceiling | 30,000 papers per query, which is arXiv's own depth limit |
| Account needed | None, and no API key |
| Price | $2.00 per 1,000 papers, flat on every plan. The free plan's $5 a month covers about 2,500 |
🔍 What arXiv Scraper does
It runs your query against arXiv and pages through the results, pausing between pages because arXiv asks callers to. A paper that turns up on two pages is delivered once.
The query uses arXiv's own syntax. It is worth two minutes, because it is what makes a search precise:
| Prefix | Searches |
|---|---|
all: | every field |
ti: | the title |
au: | the authors |
abs: | the abstract |
cat: | the category, like cs.CL or math.AG |
Combine them with AND, OR and ANDNOT. ti:attention AND au:vaswani does exactly what it looks
like. Results come back by relevance, by submission date or by last revision, and both date orders
put the newest first.
📋 What data you get from each arXiv paper
| What you get | Field |
|---|---|
| The arXiv id, without its version marker | arxivId |
| The title and the whole abstract | title, abstract |
| Every author, in the order on the paper | authors |
| The category it was filed under, and all the others | primaryCategory, categories |
| When it was first submitted and when it was last revised | publishedAt, updatedAt |
| The journal DOI, when the author added one | doi |
| The abstract page and the PDF | absUrl, pdfUrl |
▶️ How to scrape arXiv papers
- Open arXiv Scraper and click Try for free.
- Type a query into Search query, for example
cat:cs.CL AND abs:retrieval. - Pick a Sort by. Use
submittedDatewhen what you care about is new work. - Set Max papers. Start around 20 while you tune the query, then click Start.
- Download the dataset as JSON, CSV or Excel, or read it from the Apify API.
💰 How much does it cost to scrape arXiv?
$2.00 per 1,000 papers. Flat on every Apify plan, no volume tiers. On the free plan, the $5 Apify gives you each month covers about 2,500 papers.
You pay per paper row delivered. Duplicates across pages are dropped before they reach you, diagnostic rows are not charged, and a query that matches nothing costs you nothing. Set a maximum cost on the run and it stops when it gets there.
📥 What you give it
{"query": "all:large language models","sortBy": "relevance","maxItems": 50}
| Field | Default | What it is |
|---|---|---|
query | none, the form starts with all:large language models | The arXiv query. Use the prefixes above. A bare phrase with no prefix is not what arXiv expects. |
sortBy | relevance | relevance, submittedDate or lastUpdatedDate. Both date orders are newest first. |
maxItems | 50 | Papers to return, from 1 to 30,000. |
notionConnector | none | Optional. Writes every paper into your own Notion as well as the dataset, a quick start on a literature review. |
notionParentId | none | Optional. The Notion data source to write into. Leave it empty and the pages land privately in your workspace. |
proxyConfiguration | off | Optional, and off by default because arXiv is public and a normal run does not need it. |
There is no date-range field. To narrow by time, sort on submittedDate and cut the rows on
publishedAt yourself.
📤 What you get back
A real row from a run on 2 October 2026, with the abstract cut short:
{"ok": true,"charged": true,"arxivId": "2501.05032","title": "Enhancing Human-Like Responses in Large Language Models","abstract": "This paper explores the advancements in making large language models (LLMs) more human-like. ...","authors": ["Ethem Yağız Çalık", "Talha Rüzgar Akkuş"],"primaryCategory": "cs.CL","categories": ["cs.CL", "cs.AI"],"publishedAt": "2025-01-09T07:44:06Z","updatedAt": "2026-02-01T13:32:03Z","doi": null,"absUrl": "https://arxiv.org/abs/2501.05032v2","pdfUrl": "https://arxiv.org/pdf/2501.05032v2"}
| Field | How to read it |
|---|---|
arxivId | The id with any trailing version marker removed, so 2501.05032v2 arrives as 2501.05032. Stable, and the dedupe key inside a run. |
abstract | The whole abstract, with line breaks collapsed into single spaces. |
publishedAt, updatedAt | First submission and latest revision. On a paper never revised they are the same. Some rows carry the submission day only, at midnight, with updatedAt null. |
doi | null on most papers. arXiv only has one when the author added it after journal publication. |
absUrl, pdfUrl | Usually the exact version the search returned, like v2 above. Nothing is downloaded for you. |
primaryCategory | The one category the author filed it under. categories holds all of them, without repeats. |
🧾 Reading the output
Two kinds of row can land in your dataset.
| Row | How to spot it | Billed |
|---|---|---|
| A paper | ok: true, charged: true and an arxivId | yes |
| A diagnostic | ok: false and an errorCode | no |
charged: true is simply the marker every paper row carries. Use it, or ok, to tell results from
diagnostics, and do not read it as a receipt.
| Code | What it means |
|---|---|
BAD_INPUT | The query was empty. Give one in arXiv syntax. |
NO_RESULTS | The query ran and matched no papers. Loosen it, or check the prefix spelling. |
NOT_FOUND | arXiv answered 404 for the request. |
RATE_LIMITED | arXiv asked for a slower pace. Re-run with a smaller maxItems. |
SERVER_ERROR | arXiv answered with a server error. Usually passes on its own. |
BLOCKED | arXiv would not serve the search this time. |
NETWORK | arXiv could not be reached. |
Rows arrive at the end. Everything is collected before anything is written, so a large run shows an empty dataset until it finishes. Keep a first run small while you get the query right.
The default table view has no ok column, so a diagnostic row renders as one blank line. Switch
to All fields or export as JSON.
💡 What people use it for
- Watching a category like
cat:cs.CLon a schedule, sorted bysubmittedDate, and keeping only thearxivIdvalues you have not seen before. - Building a reading list on a topic, with every abstract already in the row to read or summarise.
- Following one author with
au:to see what they have put up and when it was last revised. - Collecting PDF links for a set of papers to fetch and process in your own pipeline.
🚧 What it does not do
- No full text. Abstracts and links. The PDF is yours to fetch from
pdfUrl. - No citation counts, no references, no related-paper graph. arXiv does not publish them.
- No peer-review status. A paper on arXiv may be a preprint, a submitted manuscript or a published article, and nothing on the row tells you which.
doiis usuallynull. Most preprints have none, and arXiv does not go looking.- No date-range filter, and no ascending order. Sort by date and trim the rows yourself.
- About 30,000 results per query, at most. That is arXiv's depth limit, not a setting here. A pull that size runs for close to an hour, the default run timeout, and rows are only written at the end, so split a broad query into narrower ones rather than raising the number.
- The run finishes even when every request failed, so read
okrather than the run status.
🧭 Which research scraper do you need?
| If you want | Use |
|---|---|
| arXiv preprints with full abstracts and PDF links | This one |
| Anything with a registered DOI, with its journal, publisher and ISSN | Crossref Scraper |
| Papers with readable abstracts, institutions and open-access links | OpenAlex Scraper |
| Patents rather than papers | Google Patents Search Scraper |
| Book records by keyword, title or ISBN | Books Scraper |
| Encyclopedia articles as plain text | Wikipedia Scraper |
| arXiv, OpenAlex and Wikipedia searches as tools for an AI agent | Research MCP Server |
❓ Questions people ask
Do I need an arXiv account or an API key?
No. arXiv's search is open to anyone.
Why did my arXiv search return nothing?
Usually the syntax. A bare phrase with no prefix is not a valid arXiv query; start it with all:
and try again.
Can I download the papers themselves?
Not from this actor. pdfUrl gives you a direct link to fetch.
Which date should I sort on?
submittedDate for when work first appeared, lastUpdatedDate for what was revised recently. They
differ a lot on papers that keep being updated.
Can I call it from code or connect it to an AI assistant?
Yes. The API tab has ready-made code for
Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect
https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/arxiv-scraper. Either way the run
happens on your Apify account at the same price.
Is scraping arXiv legal?
arXiv publishes its metadata for public reuse and asks callers to be polite about request rate, which this actor is. Individual papers carry their own licences, so check those before republishing text. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID and the exact query string you used. The
errorCode and hint on the diagnostic row usually name the problem on their own.