arXiv Scraper - Papers, Abstracts & PDF Links avatar

arXiv Scraper - Papers, Abstracts & PDF Links

Pricing

from $1.94 / 1,000 papers

Go to Apify Store
arXiv Scraper - Papers, Abstracts & PDF Links

arXiv Scraper - Papers, Abstracts & PDF Links

Search arXiv papers by title, author, abstract or category, in arXiv's own query syntax. Each row holds the arXiv id, title, full abstract, authors, categories, first-submitted and last-revised dates, the DOI when there is one, and abstract and PDF links. $2.00 per 1,000 papers.

Pricing

from $1.94 / 1,000 papers

Rating

5.0

(1)

Developer

Dami's Studio

Dami's Studio

Maintained by Community

Actor stats

0

Bookmarked

4

Total users

1

Monthly active users

2 days ago

Last modified

Share

Write an arXiv query and get one row per paper: the arXiv id, the title, the whole abstract, every author, the categories, the first-submitted and last-revised dates, the DOI when there is one, and direct links to the abstract page and the PDF.

Metadata and links only. The PDF link is there for you to fetch; this actor does not download or read the paper, and arXiv publishes no citation counts, so there are none here either.

InputAn arXiv search query, using arXiv's own field prefixes
OutputOne row per paper: id, title, full abstract, authors, categories, both dates, DOI, abstract and PDF links
Ceiling30,000 papers per query, which is arXiv's own depth limit
Account neededNone, and no API key
Price$2.00 per 1,000 papers, flat on every plan. The free plan's $5 a month covers about 2,500

🔍 What arXiv Scraper does

It runs your query against arXiv and pages through the results, pausing between pages because arXiv asks callers to. A paper that turns up on two pages is delivered once.

The query uses arXiv's own syntax. It is worth two minutes, because it is what makes a search precise:

PrefixSearches
all:every field
ti:the title
au:the authors
abs:the abstract
cat:the category, like cs.CL or math.AG

Combine them with AND, OR and ANDNOT. ti:attention AND au:vaswani does exactly what it looks like. Results come back by relevance, by submission date or by last revision, and both date orders put the newest first.

📋 What data you get from each arXiv paper

What you getField
The arXiv id, without its version markerarxivId
The title and the whole abstracttitle, abstract
Every author, in the order on the paperauthors
The category it was filed under, and all the othersprimaryCategory, categories
When it was first submitted and when it was last revisedpublishedAt, updatedAt
The journal DOI, when the author added onedoi
The abstract page and the PDFabsUrl, pdfUrl

▶️ How to scrape arXiv papers

  1. Open arXiv Scraper and click Try for free.
  2. Type a query into Search query, for example cat:cs.CL AND abs:retrieval.
  3. Pick a Sort by. Use submittedDate when what you care about is new work.
  4. Set Max papers. Start around 20 while you tune the query, then click Start.
  5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.

💰 How much does it cost to scrape arXiv?

$2.00 per 1,000 papers. Flat on every Apify plan, no volume tiers. On the free plan, the $5 Apify gives you each month covers about 2,500 papers.

You pay per paper row delivered. Duplicates across pages are dropped before they reach you, diagnostic rows are not charged, and a query that matches nothing costs you nothing. Set a maximum cost on the run and it stops when it gets there.

📥 What you give it

{
"query": "all:large language models",
"sortBy": "relevance",
"maxItems": 50
}
FieldDefaultWhat it is
querynone, the form starts with all:large language modelsThe arXiv query. Use the prefixes above. A bare phrase with no prefix is not what arXiv expects.
sortByrelevancerelevance, submittedDate or lastUpdatedDate. Both date orders are newest first.
maxItems50Papers to return, from 1 to 30,000.
notionConnectornoneOptional. Writes every paper into your own Notion as well as the dataset, a quick start on a literature review.
notionParentIdnoneOptional. The Notion data source to write into. Leave it empty and the pages land privately in your workspace.
proxyConfigurationoffOptional, and off by default because arXiv is public and a normal run does not need it.

There is no date-range field. To narrow by time, sort on submittedDate and cut the rows on publishedAt yourself.

📤 What you get back

A real row from a run on 2 October 2026, with the abstract cut short:

{
"ok": true,
"charged": true,
"arxivId": "2501.05032",
"title": "Enhancing Human-Like Responses in Large Language Models",
"abstract": "This paper explores the advancements in making large language models (LLMs) more human-like. ...",
"authors": ["Ethem Yağız Çalık", "Talha Rüzgar Akkuş"],
"primaryCategory": "cs.CL",
"categories": ["cs.CL", "cs.AI"],
"publishedAt": "2025-01-09T07:44:06Z",
"updatedAt": "2026-02-01T13:32:03Z",
"doi": null,
"absUrl": "https://arxiv.org/abs/2501.05032v2",
"pdfUrl": "https://arxiv.org/pdf/2501.05032v2"
}
FieldHow to read it
arxivIdThe id with any trailing version marker removed, so 2501.05032v2 arrives as 2501.05032. Stable, and the dedupe key inside a run.
abstractThe whole abstract, with line breaks collapsed into single spaces.
publishedAt, updatedAtFirst submission and latest revision. On a paper never revised they are the same. Some rows carry the submission day only, at midnight, with updatedAt null.
doinull on most papers. arXiv only has one when the author added it after journal publication.
absUrl, pdfUrlUsually the exact version the search returned, like v2 above. Nothing is downloaded for you.
primaryCategoryThe one category the author filed it under. categories holds all of them, without repeats.

🧾 Reading the output

Two kinds of row can land in your dataset.

RowHow to spot itBilled
A paperok: true, charged: true and an arxivIdyes
A diagnosticok: false and an errorCodeno

charged: true is simply the marker every paper row carries. Use it, or ok, to tell results from diagnostics, and do not read it as a receipt.

CodeWhat it means
BAD_INPUTThe query was empty. Give one in arXiv syntax.
NO_RESULTSThe query ran and matched no papers. Loosen it, or check the prefix spelling.
NOT_FOUNDarXiv answered 404 for the request.
RATE_LIMITEDarXiv asked for a slower pace. Re-run with a smaller maxItems.
SERVER_ERRORarXiv answered with a server error. Usually passes on its own.
BLOCKEDarXiv would not serve the search this time.
NETWORKarXiv could not be reached.

Rows arrive at the end. Everything is collected before anything is written, so a large run shows an empty dataset until it finishes. Keep a first run small while you get the query right.

The default table view has no ok column, so a diagnostic row renders as one blank line. Switch to All fields or export as JSON.

💡 What people use it for

  • Watching a category like cat:cs.CL on a schedule, sorted by submittedDate, and keeping only the arxivId values you have not seen before.
  • Building a reading list on a topic, with every abstract already in the row to read or summarise.
  • Following one author with au: to see what they have put up and when it was last revised.
  • Collecting PDF links for a set of papers to fetch and process in your own pipeline.

🚧 What it does not do

  • No full text. Abstracts and links. The PDF is yours to fetch from pdfUrl.
  • No citation counts, no references, no related-paper graph. arXiv does not publish them.
  • No peer-review status. A paper on arXiv may be a preprint, a submitted manuscript or a published article, and nothing on the row tells you which.
  • doi is usually null. Most preprints have none, and arXiv does not go looking.
  • No date-range filter, and no ascending order. Sort by date and trim the rows yourself.
  • About 30,000 results per query, at most. That is arXiv's depth limit, not a setting here. A pull that size runs for close to an hour, the default run timeout, and rows are only written at the end, so split a broad query into narrower ones rather than raising the number.
  • The run finishes even when every request failed, so read ok rather than the run status.

🧭 Which research scraper do you need?

If you wantUse
arXiv preprints with full abstracts and PDF linksThis one
Anything with a registered DOI, with its journal, publisher and ISSNCrossref Scraper
Papers with readable abstracts, institutions and open-access linksOpenAlex Scraper
Patents rather than papersGoogle Patents Search Scraper
Book records by keyword, title or ISBNBooks Scraper
Encyclopedia articles as plain textWikipedia Scraper
arXiv, OpenAlex and Wikipedia searches as tools for an AI agentResearch MCP Server

❓ Questions people ask

Do I need an arXiv account or an API key?

No. arXiv's search is open to anyone.

Why did my arXiv search return nothing?

Usually the syntax. A bare phrase with no prefix is not a valid arXiv query; start it with all: and try again.

Can I download the papers themselves?

Not from this actor. pdfUrl gives you a direct link to fetch.

Which date should I sort on?

submittedDate for when work first appeared, lastUpdatedDate for what was revised recently. They differ a lot on papers that keep being updated.

Can I call it from code or connect it to an AI assistant?

Yes. The API tab has ready-made code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/arxiv-scraper. Either way the run happens on your Apify account at the same price.

arXiv publishes its metadata for public reuse and asks callers to be polite about request rate, which this actor is. Individual papers carry their own licences, so check those before republishing text. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the run ID and the exact query string you used. The errorCode and hint on the diagnostic row usually name the problem on their own.