PubMed Scraper: Articles, Abstracts & Citations avatar

PubMed Scraper: Articles, Abstracts & Citations

Pricing

from $12.00 / 1,000 article records

Go to Apify Store
PubMed Scraper: Articles, Abstracts & Citations

PubMed Scraper: Articles, Abstracts & Citations

Scrape PubMed articles by keyword, PubMed syntax or PMID: title, journal, authors, DOI, abstract, MeSH terms and keywords, plus cited-by and related papers as rows. Filter by date, sort by relevance. Export CSV, Excel, JSON or XML. No login or API key needed.

Pricing

from $12.00 / 1,000 article records

Rating

0.0

(0)

Developer

RecordsData

RecordsData

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

18 hours ago

Last modified

Categories

Share

PunkRecordsData

๐Ÿงซ PubMed Scraper: Articles, Abstracts & Citations

PubMed Scraper exports biomedical articles from PubMed for any search query or list of PMIDs: title, journal, authors, DOI, PMCID, publication types, full abstract, MeSH terms and keywords, plus cited-by and related articles as extra rows. It uses NCBI's official E-utilities, needs no login and no API key, and exports CSV, Excel, JSON or XML. Pay per result, nothing for empty or failed rows.

The PubMed Scraper turns a PubMed query into a clean dataset. A search for "crispr gene editing" matched 25,537 articles on PubMed when this page was written, and a 10-article cloud run delivered 10 rows with title, journal and DOI, 9 of them with a structured abstract. It is built for literature reviews, pharma monitoring, bibliometrics and medical AI pipelines that need abstracts and MeSH indexing without copy-pasting.

๐Ÿ”Ž What does the PubMed Scraper do?

  • Search PubMed by keyword: full PubMed syntax, including boolean operators and field tags such as smith j[au], [ti] and [mh].
  • Fetch exact articles by PMID: paste a list of PubMed IDs and get one row each.
  • Pull abstracts and MeSH terms: the full structured abstract, MeSH headings and author keywords per article.
  • Collect cited-by articles: papers that cite each result, one full bibliographic row per citing paper.
  • Collect related articles: PubMed's similar-articles list, one row per related paper.
  • Filter by publication date: after and before dates in YYYY/MM/DD, sorted by relevance or newest first.

๐Ÿ“‹ What data can you extract from PubMed?

FieldDescription
recordTypearticle, cited-by or related
pmid, urlPubMed ID and link to the PubMed page
title, journalArticle title and full journal name
authorsUp to 15 authors per article
pubDate, epubDatePublication and electronic publication dates
volume, issue, pagesCitation details
doi, pmcidDOI and PubMed Central ID when a free full text exists
publicationTypes, languageFor example Journal Article, Review; language code
abstractFull structured abstract with section labels
meshTerms, keywordsMeSH headings and author keywords
sourcePmid, sourceTitleOn cited-by and related rows: the article they came from
searchTerm, scrapedAtThe query that found the row and the UTC timestamp

Optional fields that PubMed does not provide for an article (for example issue or pmcid) hold the text N/A or Not Disclosed, so the column layout stays the same across rows.

๐Ÿ“Š Sample output of the PubMed Scraper

Real record from a cloud run with the query crispr gene editing (abstract shortened here):

{
"recordType": "article",
"pmid": "31295471",
"title": "CRISPR-Cas9 system: A new-fangled dawn in gene editing.",
"url": "https://pubmed.ncbi.nlm.nih.gov/31295471/",
"journal": "Life sciences",
"authors": ["Gupta D", "Bhattacharjee O", "Mandal D", "Sen MK"],
"pubDate": "2019 Sep 1",
"epubDate": "2019 Jul 8",
"volume": "232",
"issue": "N/A",
"pages": "116636",
"doi": "10.1016/j.lfs.2019.116636",
"pmcid": "N/A",
"publicationTypes": ["Journal Article", "Review"],
"language": "eng",
"searchTerm": "crispr gene editing",
"abstract": "Till date, only three techniques namely Zinc Finger Nuclease (ZFN), Transcription-Activator Like Effector Nucleases (TALEN) and Clustered Regularly Interspaced Short Palindromic Repeats-CRISPR-Associa...",
"meshTerms": ["Animals", "CRISPR-Cas Systems", "Gene Editing", "Genome", "Humans", "Plants"],
"keywords": ["CRISPR-Cas9", "Genome editing", "Knock in", "Knock out"],
"scrapedAt": "2026-10-04T05:04:00.654Z"
}

๐Ÿ’ต How much does it cost to scrape PubMed?

The actor is pay per event. You are charged only for rows that were saved with real data. Rows for PMIDs that do not exist, articles without a title, and empty searches are never charged.

EventWhat you getPrice (free tier)
Article recordOne article with journal, authors, dates, DOI and publication types$15.00 per 1,000
Abstract + MeSHFull abstract, MeSH headings and keywords, only when the article has an abstract$12.00 per 1,000
Cited-by articleOne article citing a result, as a full row$10.00 per 1,000
Related articleOne related article from PubMed similar-articles$10.00 per 1,000

Prices are the live prices at the time of writing and drop on higher Apify plans. 1,000 articles with abstracts cost about $27 at the free tier. With abstracts switched off (includeAbstract: false) they cost $15. Free users get a 10-article preview per run. You can set a maximum charge per run in Apify and the actor stops cleanly when it is reached.

๐Ÿš€ How to scrape PubMed (3 steps)

  1. Open the actor and click Try for free.
  2. Enter search terms (or PMIDs), choose sorting, an optional date range and which modules you want: abstracts, cited-by, related.
  3. Click Start, then download the dataset as CSV, Excel, JSON or XML.

โš™๏ธ Input of the PubMed Scraper

FieldTypeMeaning
searchTermslist of textPubMed queries, one search per term
pmidslist of textExact PubMed IDs to fetch
maxItemsnumberMax articles to return (free users: 10)
sortByrelevance or pub_dateBest match or most recent first
publishedAfter, publishedBeforetext YYYY/MM/DDPublication date range
includeAbstractboolean (default true)Abstract, MeSH headings and keywords
includeCitedBy, maxCitedByPerArticleboolean, numberCiting papers, capped per article (default 20)
includeRelated, maxRelatedPerArticleboolean, numberRelated papers, capped per article (default 10)

Provide at least one search term or one PMID. If both are empty the run fails immediately with a clear message and nothing is charged.

{
"searchTerms": ["crispr gene editing"],
"maxItems": 10,
"sortBy": "relevance",
"includeAbstract": true,
"includeCitedBy": false,
"includeRelated": false
}

๐Ÿ“ฆ Output of the PubMed Scraper

One dataset row per article, cited-by paper or related paper, with recordType telling them apart. The Overview view shows type, PMID, title, journal, date, DOI and link. The dataset downloads as JSON, CSV, Excel or XML. A search with no matches ends as succeeded with the status message "no items matched" and charges nothing. If PubMed itself is unreachable and no article was delivered, the run fails instead of pretending to be empty.

๐Ÿ“ˆ PubMed scraper vs alternatives

Prices measured through the public Apify Store API on 2026-10-03 (free tier, per 1,000 results):

ActorPrice per 1,000Notes
This actor$15 (article) + $12 (abstract)Abstracts, MeSH, cited-by and related as separate events
scrapestorm/pubmed-articles-scraper$2.89One flat event
easyapi/pubmed-search-scraper$2.99 plus $0.09 per startOne flat event
ryanclinton/pubmed-research-search$2.00One flat event
labrat011/pubmed-scraper$0.80One flat event

We are more expensive per row than the flat-price actors. What the extra buys: structured abstracts with MeSH headings and keywords in the same run, and citation snowballing (cited-by and related rows) without a second tool. If you only need bibliographic basics, a cheaper actor is the better pick.

๐Ÿ’ผ Use cases

  • Systematic literature reviews: export a full query with abstracts into a screening sheet.
  • Pharma and competitor monitoring: scheduled weekly runs on a molecule or indication, newest first.
  • Bibliometrics: citation snowballing using cited-by rows.
  • Medical AI and RAG: abstracts plus MeSH terms as clean retrieval data.

๐Ÿ”Œ Run via API, schedule and integrations

curl -X POST "https://api.apify.com/v2/acts/recordsdata~pubmed-articles-scraper/runs?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"searchTerms":["breast cancer AND immunotherapy"],"maxItems":50}'

Use the Apify client for Python or JavaScript, schedule runs, or connect results to Zapier, Make, n8n, Google Sheets or Slack through Apify integrations. The actor is also available to AI agents through the Apify MCP server.

The actor reads public bibliographic data through NCBI's official E-utilities, at a pace under the keyless limit of 3 requests per second, and identifies itself in each request. It does not log in or bypass anything. You are responsible for how you use the data and for respecting NCBI and publisher terms. Abstracts remain the property of their publishers.

โ“ Frequently asked questions

๐Ÿงซ How do I export PubMed search results to CSV or Excel?

Enter your query, click Start and download the dataset from the Storage tab as CSV, Excel, JSON or XML.

๐Ÿ“„ Does it include full abstracts?

Yes. With includeAbstract on, each article gets its structured abstract, MeSH headings and keywords. The abstract event is charged only when an abstract exists. Otherwise the field says "Not Available".

๐Ÿ”Ž Can I use PubMed advanced query syntax?

Yes. Boolean operators, phrase quotes and field tags like [au], [ti] and [mh] work as on pubmed.ncbi.nlm.nih.gov.

๐Ÿ•ธ How do I get papers that cite an article?

Turn on includeCitedBy. Each citing paper becomes a full row with recordType set to cited-by and the source PMID attached.

๐Ÿ†” Can I fetch specific articles by PMID?

Yes. Put the IDs in pmids. They run before the searches. A PMID that does not exist returns a free error row.

๐Ÿ“… Can I filter by publication date?

Yes, with publishedAfter and publishedBefore in YYYY/MM/DD format.

๐Ÿ’ต Do I pay for empty or failed rows?

No. Not-found PMIDs, articles without a title and searches with no results are not charged. If the run hits your maximum charge, it stops cleanly.

๐Ÿ”ข How many articles can one run return?

Up to 1,000,000 on paid plans, but PubMed limits a single query to 9,999 records, so split large topics by date range. Free users get 10 articles per run.

โš™๏ธ Does it need an NCBI API key?

No.

๐Ÿšซ Why did I get 0 results?

Your query matched nothing on PubMed (try fewer terms or remove the date filter), or you left both searchTerms and pmids empty, which fails the run. Check the run status message.

๐Ÿ”— Want more research and health data? Other PunkRecordsData scrapers

๐Ÿ†˜ Support

Questions or feature requests: use the Issues tab on this actor page or write to contact.punkrecordsdata@gmail.com. Independent tool, not affiliated with NCBI, NLM or NIH. Only public data. Not medical advice.

Last updated: 2026-10-03