arXiv Preprint Scraper
Pricing
from $8.00 / 1,000 results
arXiv Preprint Scraper
Scrapes arXiv preprints by search query, category, or paper ID and returns each paper's metadata as a flat row. Supports bulk collection up to a million papers per run.
Pricing
from $8.00 / 1,000 results
Rating
5.0
(1)
Developer
ParseForge
Maintained by CommunityActor stats
1
Bookmarked
20
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
arXiv Preprint Scraper
Scrape arXiv preprints by keyword, category, or paper ID. Every paper comes with its title, authors, full abstract, categories, submission and revision dates, DOI, journal reference, and PDF link. No API key or registration. Export to CSV, JSON, Excel, or XML.
arXiv's official API rate-limits you and returns raw Atom feeds you have to parse yourself. This Actor reads the public search and abstract pages directly, filters by category or sort order, and returns each matching paper in one clean, flat row. It handles the pagination and parsing so you get structured data from a simple search query or a list of arXiv IDs.
| Who uses it | What they scrape arXiv for |
|---|---|
| Academic researchers | Building a systematic literature review dataset for a specific topic |
| Data scientists | Tracking new papers in a niche ML subfield each week |
| Librarians | Monitoring institutional author publications across multiple categories |
| Journalists | Finding the latest preprints on a breaking science story |
What it does
This Actor collects arXiv preprints by search query, category, direct URL, or paper ID, and returns each one as a flat row with its metadata.
- π Keyword search: searches titles, abstracts, and author names for your terms.
- π Category filter: restrict results to a single arXiv category like cs.LG or physics.optics.
- π Direct paper fetch: supply a list of arXiv IDs or a single paper URL to get exactly those records.
- π Date sorting: order by submission date or last-updated date, ascending or descending.
Results export to CSV, JSON, Excel, or XML, or straight from the API.
What you can do with arXiv data
π Build a literature review dataset.
A PhD student searches for "retrieval augmented generation" in cs.CL, collects 500 papers, and exports them to CSV for screening in a systematic review.
π Monitor a research field weekly.
A data scientist runs the Actor every Monday with a category filter and sort by submission date to catch the latest preprints in stat.ML.
π Fetch specific papers by ID.
A research lab manager pastes a list of 200 arXiv IDs from a conference proceedings page and gets back full metadata for their internal database.
π° Gather sources for a science story.
A journalist searches for a recent discovery keyword, sorts by relevance, and pulls the top 20 preprints to cite in an article.
Why choose this scraper
| What you get | |
|---|---|
| No API key needed | Reads public pages, no registration or OAuth. |
| Flat, predictable schema | Every paper returns the same fields, ready for analysis. |
| Bulk collection | Up to 10,000 papers per search query, paginated for you. Split larger pulls by category or query. |
| Multiple input modes | Search by keyword, browse a category, or fetch specific IDs. |
How it compares
This Actor focuses on direct arXiv scraping with multiple input modes and bulk collection, while the competitors below offer MCP server integrations or multi-source search.
| Feature | ParseForge | Academic Research MCP | Academic Paper Scraper |
|---|---|---|---|
| Keyword search on arXiv | Yes | Yes | Yes |
| Fetch by arXiv ID list | Yes | Not listed | Not listed |
| Category filter | Yes | Not listed | Not listed |
| Sort by submission date | Yes | Not listed | Not listed |
| MCP server for AI agents | Not listed | Yes | Yes |
| Google Scholar search | Not listed | Yes | Not listed |
| Semantic Scholar search | Not listed | Not listed | Yes |
Configure the run
Drive the Actor from a search query, an arXiv category, a list of paper IDs, or a direct URL, alone or together, and the sort order and maximum items filter the results before they reach your dataset. The Input tab lists every parameter.
A first run with the defaults:
{"searchQuery": "large language models","maxItems": 25}
A larger pull:
{"searchQuery": "large language models","maxItems": 200}
Output
Each paper is one row:
{"arxivId": "1706.03762","version": "v7","title": "Attention Is All You Need","authors": ["Ashish Vaswani", "Noam Shazeer", "Niki Parmar", "Jakob Uszkoreit", "Llion Jones", "Aidan N. Gomez", "Lukasz Kaiser", "Illia Polosukhin"],"summary": "The dominant sequence transduction models are based on complex recurrent or convolutional neural networks...","primaryCategory": "cs.CL","categories": ["cs.CL", "cs.LG"],"published": "2017-06-12","updated": "2023-08-02","doi": null,"journalRef": null,"comments": "15 pages, 5 figures","absUrl": "https://arxiv.org/abs/1706.03762v7","pdfUrl": "https://arxiv.org/pdf/1706.03762v7","scrapedAt": "2026-09-14T14:14:29.594Z"}
| Field | What it holds |
|---|---|
arxivId, version | arXiv identifier and the latest version number |
title, authors, summary | Title, author names in order, and the full abstract |
primaryCategory, categories | Primary arXiv category and every cross-listed category |
published, updated | Date of the first version and of the latest revision |
doi, journalRef | Journal DOI and journal reference, when the authors added them |
comments | Author comments such as page count or conference |
absUrl, pdfUrl | Abstract page and PDF of the latest version |
Pricing
Pay per event: a small fee when the run starts, plus a fee for each paper saved to your dataset. Runs that find nothing pay no per-paper fee.
| Apify plan | Run start | Per paper | 1,000 papers |
|---|---|---|---|
| Free (10-paper preview) | $0.16 | $0.012 | n/a |
| Starter | $0.123 | $0.0107 | about $10.79 |
| Scale | $0.087 | $0.0093 | about $9.42 |
| Business and above | $0.05 | $0.008 | about $8.05 |
The start fee is charged once per GB of run memory; the default 128 MB run pays it once. New Apify accounts start with $5 in free credit.
Free users
Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to lift the preview limit.
Run it
- Create a free Apify account with $5 in credit.
- Open the arXiv Preprint Scraper.
- Set your inputs and any filters, then click Start.
- Export the results as CSV, Excel, JSON, or XML from the Dataset tab.
Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.
Use with AI agents (MCP)
Give an AI agent live access to arXiv through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:
$claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/arxiv-scraper"
Then prompt it in plain language to run the scraper and read back the results.
Troubleshooting
Why am I getting no results?
Check that your search query is spelled correctly and that the category, if set, is a valid arXiv category code like cs.AI or physics.optics. Also try removing the category filter to see if results appear without it.
The Actor is running very slowly.
Large maxItems values mean the Actor must paginate through many pages. Reduce the maximum papers or narrow your search with a more specific query or a category filter to speed it up.
I got fewer papers than my maxItems setting.
The Actor stops when arXiv has no more results matching your query. Try a broader search term or remove the category filter to find more papers.
Some fields are empty in my output.
arXiv does not guarantee every field for every paper. For example, some older preprints may lack a DOI or a specific author affiliation. Empty fields mean the data was not present on the page.
The sort order does not seem to work.
The sortOrder field only applies when sortBy is set to submittedDate or lastUpdatedDate. When sortBy is relevance, the order is fixed by arXiv's search engine.
FAQ
| Question | Answer |
|---|---|
| Do I need an arXiv API key? | No. This Actor reads the public arXiv website directly. You do not need to register an application or manage API credentials. |
| How many papers can I scrape in one run? | arXiv's search returns at most 10,000 papers per query, and the Actor paginates up to that point or your maxItems. Past 10,000 it tries the arXiv API, which is often rate-limited, so for bigger pulls run several narrower queries or one per category. |
| Can I scrape a specific arXiv category without a search term? | Yes. Leave the search query empty and set a category like cs.LG. The Actor will return the most recent papers in that category, sorted by your chosen order. |
| What output formats are supported? | You can export your dataset as CSV, JSON, Excel, or XML from the Apify platform after the run finishes. |
| Does this Actor get the full PDF text? | No. It collects the metadata shown on the abstract page: title, authors, abstract, submission date, and related fields. It does not download or parse the full PDF. |
| Can I scrape papers by a specific author? | Yes. Enter the author's name in the search query field. arXiv's search engine matches against author names, so a query like "Yann LeCun" will return papers by that author. |
| What happens if I provide both a search query and a list of arXiv IDs? | The arXiv IDs field takes priority. When you supply paper IDs, the Actor fetches those specific papers and ignores the search query and category. |
| Is there a rate limit? | arXiv's official API is frequently rate-limited for everyone. This Actor reads arXiv's public search and abstract pages, waits between requests, and retries within the run's time limit, falling back to the API only when the website does not answer. |
| Can I schedule this Actor to run automatically? | Yes. Apify supports scheduled runs. You can set it to run daily or weekly to collect new papers as they appear on arXiv. |
Related actors
Browse the full ParseForge collection for more scrapers.
π Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.
β οΈ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by arXiv, Cornell University. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.
