Semantic Scholar Scraper avatar

Semantic Scholar Scraper

Pricing

from $8.00 / 1,000 results

Go to Apify Store
Semantic Scholar Scraper

Semantic Scholar Scraper

Scrapes academic papers from Semantic Scholar search results. Returns each paper as a flat row with title, authors, year, citations, venue, and abstract. Supports year and PDF filters.

Pricing

from $8.00 / 1,000 results

Rating

1.1

(2)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

2

Bookmarked

51

Total users

3

Monthly active users

6 days ago

Last modified

Share

ParseForge

Semantic Scholar Scraper

Search Semantic Scholar for papers by title or keywords, look up authors with their h-index and papers, or fetch papers by DOI, arXiv id or Semantic Scholar id. Every paper comes with its title, authors, year, venue, journal, abstract, citation counts, fields of study, DOI/arXiv/PubMed ids and the open-access PDF link. No API key required. Export to CSV, JSON, Excel, or XML.

This Actor drives Semantic Scholar's official Graph API — clean JSON, no browser, no proxy. A title search such as "attention is all you need" returns the paper you mean first, because the query is matched as a phrase and ranked by citations; a keyword search such as "graph neural networks drug discovery" returns the most-cited matches with server-side filters for year, field of study, publication type, venue, minimum citations and PDF availability.

Who uses itWhat they scrape Semantic Scholar for
Academic researchersBuilding a corpus of papers for a systematic literature review.
PhD studentsRecent publications in a field, filtered by year and venue, to find research gaps.
Data scientistsAuthor profiles (h-index, citation totals) and paper metadata for bibliometrics.
Librarians and research officesPublication lists per author, with DOIs and open-access links.
AI agentsResolving a paper from its title, DOI or arXiv id in one call.

What it does

Three search modes, one flat row per paper or author:

  • Papers: a title or keyword query. Phrase match first (the exact paper for a title query), most-cited first by default, or newest/oldest first. Filters: year or year range, fields of study, publication types, venue, minimum citations, open-access PDF only.
  • Authors: an author name. Each author row carries affiliations, homepage, paper count, citation count, h-index, ORCID and DBLP ids. Optionally list each author's papers (newest first) after the author row.
  • Paper IDs: a list of ids — 40-character Semantic Scholar ids, DOI:10.…, ArXiv:1706.03762, PMID:…, PMCID:…, CorpusId:…. Plain DOIs and arXiv numbers are recognised without a prefix.

A semanticscholar.org URL works too: a search URL (query, year range and PDF filter are read from it), a paper URL or an author URL.

Every paper row contains: paperId, corpusId, title, authors, authorIds, firstAuthor, year, venue, journalName, journalVolume, journalPages, publicationDate, publicationTypes, abstract, url, citationCount, referenceCount, influentialCitationCount, fieldsOfStudy, s2FieldsOfStudy, doi, arxivId, pubmedId, isOpenAccess, openAccessPdfUrl, source, query, scrapedAt.

Every author row contains: authorId, name, affiliations, homepage, paperCount, citationCount, hIndex, orcid, dblp, url, query, scrapedAt.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

What you can do with Semantic Scholar data

Find the paper behind a title.

An agent sends "deep residual learning for image recognition" and gets He et al. 2015 back as the first row, with its DOI, arXiv id and 237,000 citations.

Build a literature review corpus.

A PhD candidate searches "graph neural networks drug discovery" with year 2022-, field of study Computer Science and at least 20 citations, and exports 500 abstracts to screen for relevance.

Profile an author.

A research office pulls "Yoshua Bengio" with Include each author's papers on and gets the author row (813 papers, h-index 212) followed by the publication list, newest first.

Resolve a reading list.

A lab manager pastes 200 DOIs and arXiv ids and gets one row per paper with the open-access PDF link where one exists.

Why choose this scraper

What you get
Title search that worksPhrase match plus citation ranking returns the well-known paper first instead of thousands of loosely related ones.
Three modesPapers, authors and direct id lookup in one Actor.
Server-side filtersYear, field of study, publication type, venue, minimum citations and PDF availability narrow the result count rather than costing rows.
No API key requiredRuns on the public pool at about one request per second; add a free key for relevance ranking and higher limits.
Scales to a million itemsPaid users can pull entire research fields in one run.

How it compares

FeatureParseForgeSemantic Scholar Search Scraper
Title search returns the exact paper firstYesNot listed
Author search with h-index and papersYesNot listed
Lookup by DOI, arXiv, PubMed or Corpus idYesNot listed
Year, field of study, publication type, venue and citation filtersYesNot listed
Start from a full Semantic Scholar URLYesNot listed
Optional API key for higher rate limitsYesNot listed

Configure the run

Pick the mode, give a query (or a Semantic Scholar URL, or a list of ids) and set the maximum rows. The Input tab lists every filter.

A first run with the defaults:

{
"searchType": "papers",
"query": "attention is all you need",
"maxItems": 10
}

A filtered keyword pull:

{
"searchType": "papers",
"query": "graph neural networks drug discovery",
"year": "2022-",
"fieldsOfStudy": ["Computer Science"],
"minCitationCount": 20,
"sort": "newest",
"maxItems": 200
}

An author with their papers:

{
"searchType": "authors",
"query": "Yoshua Bengio",
"includeAuthorPapers": true,
"maxItems": 100
}

Papers by id:

{
"searchType": "ids",
"ids": ["ArXiv:1706.03762", "DOI:10.1145/3065386", "204e3073870fae3d05bcbc2f6a8e263d9b72e776"]
}

Limits

  • Without an API key, Semantic Scholar's public pool allows roughly one request per second, and its relevance-ranked search endpoint is closed to anonymous traffic. The Actor therefore matches your query as a phrase and ranks by citations, which surfaces the paper you mean for a title search. A free key from semanticscholar.org/product/api switches Relevance to true relevance ranking and lifts the rate limit.
  • When the public pool is saturated the Actor backs off and retries; if it is still refused, the run writes one diagnostic row (type: "error"), which is not charged.
  • Paper ids that Semantic Scholar does not know return a diagnostic row instead of a paper row.

Pricing

Pay-per-event: $0.16 per run start plus $12 per 1,000 results (result-item, one per paper or author row). Diagnostic rows are free.

Results collectedApproximate cost
100 results$1.36
1,000 results$12.16
10,000 results$120.16

New Apify accounts start with $5 in free credit.

Free users

Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.

Run it

  1. Create a free Apify account with $5 in credit.
  2. Open the Semantic Scholar Scraper.
  3. Pick a mode, enter a query and any filters, then click Start.
  4. Export the results as CSV, Excel, JSON, or XML from the Dataset tab.

Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.

Use with AI agents (MCP)

Give an AI agent live access to Semantic Scholar through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

$claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/semantic-scholar-scraper"

Then ask: "Find the paper 'Attention is All You Need' and give me its DOI and citation count."

FAQ

Why did the old version return unrelated papers for a title? The public bulk endpoint orders results by internal id unless told otherwise, so a five-word title matched thousands of papers containing those words. The Actor now matches the words as a phrase and ranks by citations, and falls back to the bare keywords only when the phrase matches nothing.

Can I get relevance ranking like the website? Yes, with a free Semantic Scholar API key in the apiKey input.

Is the data public? Yes. Semantic Scholar publishes this metadata through its open Graph API; the Actor reads only public records.