PubMed Scraper: Articles, Abstracts & Citations
Pricing
from $12.00 / 1,000 article records
PubMed Scraper: Articles, Abstracts & Citations
Scrape PubMed articles by keyword, PubMed syntax or PMID: title, journal, authors, DOI, abstract, MeSH terms and keywords, plus cited-by and related papers as rows. Filter by date, sort by relevance. Export CSV, Excel, JSON or XML. No login or API key needed.
Pricing
from $12.00 / 1,000 article records
Rating
0.0
(0)
Developer
RecordsData
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
18 hours ago
Last modified
Categories
Share
๐งซ PubMed Scraper: Articles, Abstracts & Citations
PubMed Scraper exports biomedical articles from PubMed for any search query or list of PMIDs: title, journal, authors, DOI, PMCID, publication types, full abstract, MeSH terms and keywords, plus cited-by and related articles as extra rows. It uses NCBI's official E-utilities, needs no login and no API key, and exports CSV, Excel, JSON or XML. Pay per result, nothing for empty or failed rows.
The PubMed Scraper turns a PubMed query into a clean dataset. A search for "crispr gene editing" matched 25,537 articles on PubMed when this page was written, and a 10-article cloud run delivered 10 rows with title, journal and DOI, 9 of them with a structured abstract. It is built for literature reviews, pharma monitoring, bibliometrics and medical AI pipelines that need abstracts and MeSH indexing without copy-pasting.
๐ What does the PubMed Scraper do?
- Search PubMed by keyword: full PubMed syntax, including boolean operators and field tags such as
smith j[au],[ti]and[mh]. - Fetch exact articles by PMID: paste a list of PubMed IDs and get one row each.
- Pull abstracts and MeSH terms: the full structured abstract, MeSH headings and author keywords per article.
- Collect cited-by articles: papers that cite each result, one full bibliographic row per citing paper.
- Collect related articles: PubMed's similar-articles list, one row per related paper.
- Filter by publication date: after and before dates in
YYYY/MM/DD, sorted by relevance or newest first.
๐ What data can you extract from PubMed?
| Field | Description |
|---|---|
recordType | article, cited-by or related |
pmid, url | PubMed ID and link to the PubMed page |
title, journal | Article title and full journal name |
authors | Up to 15 authors per article |
pubDate, epubDate | Publication and electronic publication dates |
volume, issue, pages | Citation details |
doi, pmcid | DOI and PubMed Central ID when a free full text exists |
publicationTypes, language | For example Journal Article, Review; language code |
abstract | Full structured abstract with section labels |
meshTerms, keywords | MeSH headings and author keywords |
sourcePmid, sourceTitle | On cited-by and related rows: the article they came from |
searchTerm, scrapedAt | The query that found the row and the UTC timestamp |
Optional fields that PubMed does not provide for an article (for example issue or pmcid) hold the text N/A or Not Disclosed, so the column layout stays the same across rows.
๐ Sample output of the PubMed Scraper
Real record from a cloud run with the query crispr gene editing (abstract shortened here):
{"recordType": "article","pmid": "31295471","title": "CRISPR-Cas9 system: A new-fangled dawn in gene editing.","url": "https://pubmed.ncbi.nlm.nih.gov/31295471/","journal": "Life sciences","authors": ["Gupta D", "Bhattacharjee O", "Mandal D", "Sen MK"],"pubDate": "2019 Sep 1","epubDate": "2019 Jul 8","volume": "232","issue": "N/A","pages": "116636","doi": "10.1016/j.lfs.2019.116636","pmcid": "N/A","publicationTypes": ["Journal Article", "Review"],"language": "eng","searchTerm": "crispr gene editing","abstract": "Till date, only three techniques namely Zinc Finger Nuclease (ZFN), Transcription-Activator Like Effector Nucleases (TALEN) and Clustered Regularly Interspaced Short Palindromic Repeats-CRISPR-Associa...","meshTerms": ["Animals", "CRISPR-Cas Systems", "Gene Editing", "Genome", "Humans", "Plants"],"keywords": ["CRISPR-Cas9", "Genome editing", "Knock in", "Knock out"],"scrapedAt": "2026-10-04T05:04:00.654Z"}
๐ต How much does it cost to scrape PubMed?
The actor is pay per event. You are charged only for rows that were saved with real data. Rows for PMIDs that do not exist, articles without a title, and empty searches are never charged.
| Event | What you get | Price (free tier) |
|---|---|---|
| Article record | One article with journal, authors, dates, DOI and publication types | $15.00 per 1,000 |
| Abstract + MeSH | Full abstract, MeSH headings and keywords, only when the article has an abstract | $12.00 per 1,000 |
| Cited-by article | One article citing a result, as a full row | $10.00 per 1,000 |
| Related article | One related article from PubMed similar-articles | $10.00 per 1,000 |
Prices are the live prices at the time of writing and drop on higher Apify plans. 1,000 articles with abstracts cost about $27 at the free tier. With abstracts switched off (includeAbstract: false) they cost $15. Free users get a 10-article preview per run. You can set a maximum charge per run in Apify and the actor stops cleanly when it is reached.
๐ How to scrape PubMed (3 steps)
- Open the actor and click Try for free.
- Enter search terms (or PMIDs), choose sorting, an optional date range and which modules you want: abstracts, cited-by, related.
- Click Start, then download the dataset as CSV, Excel, JSON or XML.
โ๏ธ Input of the PubMed Scraper
| Field | Type | Meaning |
|---|---|---|
searchTerms | list of text | PubMed queries, one search per term |
pmids | list of text | Exact PubMed IDs to fetch |
maxItems | number | Max articles to return (free users: 10) |
sortBy | relevance or pub_date | Best match or most recent first |
publishedAfter, publishedBefore | text YYYY/MM/DD | Publication date range |
includeAbstract | boolean (default true) | Abstract, MeSH headings and keywords |
includeCitedBy, maxCitedByPerArticle | boolean, number | Citing papers, capped per article (default 20) |
includeRelated, maxRelatedPerArticle | boolean, number | Related papers, capped per article (default 10) |
Provide at least one search term or one PMID. If both are empty the run fails immediately with a clear message and nothing is charged.
{"searchTerms": ["crispr gene editing"],"maxItems": 10,"sortBy": "relevance","includeAbstract": true,"includeCitedBy": false,"includeRelated": false}
๐ฆ Output of the PubMed Scraper
One dataset row per article, cited-by paper or related paper, with recordType telling them apart. The Overview view shows type, PMID, title, journal, date, DOI and link. The dataset downloads as JSON, CSV, Excel or XML. A search with no matches ends as succeeded with the status message "no items matched" and charges nothing. If PubMed itself is unreachable and no article was delivered, the run fails instead of pretending to be empty.
๐ PubMed scraper vs alternatives
Prices measured through the public Apify Store API on 2026-10-03 (free tier, per 1,000 results):
| Actor | Price per 1,000 | Notes |
|---|---|---|
| This actor | $15 (article) + $12 (abstract) | Abstracts, MeSH, cited-by and related as separate events |
| scrapestorm/pubmed-articles-scraper | $2.89 | One flat event |
| easyapi/pubmed-search-scraper | $2.99 plus $0.09 per start | One flat event |
| ryanclinton/pubmed-research-search | $2.00 | One flat event |
| labrat011/pubmed-scraper | $0.80 | One flat event |
We are more expensive per row than the flat-price actors. What the extra buys: structured abstracts with MeSH headings and keywords in the same run, and citation snowballing (cited-by and related rows) without a second tool. If you only need bibliographic basics, a cheaper actor is the better pick.
๐ผ Use cases
- Systematic literature reviews: export a full query with abstracts into a screening sheet.
- Pharma and competitor monitoring: scheduled weekly runs on a molecule or indication, newest first.
- Bibliometrics: citation snowballing using cited-by rows.
- Medical AI and RAG: abstracts plus MeSH terms as clean retrieval data.
๐ Run via API, schedule and integrations
curl -X POST "https://api.apify.com/v2/acts/recordsdata~pubmed-articles-scraper/runs?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"searchTerms":["breast cancer AND immunotherapy"],"maxItems":50}'
Use the Apify client for Python or JavaScript, schedule runs, or connect results to Zapier, Make, n8n, Google Sheets or Slack through Apify integrations. The actor is also available to AI agents through the Apify MCP server.
โ๏ธ Is it legal to scrape PubMed?
The actor reads public bibliographic data through NCBI's official E-utilities, at a pace under the keyless limit of 3 requests per second, and identifies itself in each request. It does not log in or bypass anything. You are responsible for how you use the data and for respecting NCBI and publisher terms. Abstracts remain the property of their publishers.
โ Frequently asked questions
๐งซ How do I export PubMed search results to CSV or Excel?
Enter your query, click Start and download the dataset from the Storage tab as CSV, Excel, JSON or XML.
๐ Does it include full abstracts?
Yes. With includeAbstract on, each article gets its structured abstract, MeSH headings and keywords. The abstract event is charged only when an abstract exists. Otherwise the field says "Not Available".
๐ Can I use PubMed advanced query syntax?
Yes. Boolean operators, phrase quotes and field tags like [au], [ti] and [mh] work as on pubmed.ncbi.nlm.nih.gov.
๐ธ How do I get papers that cite an article?
Turn on includeCitedBy. Each citing paper becomes a full row with recordType set to cited-by and the source PMID attached.
๐ Can I fetch specific articles by PMID?
Yes. Put the IDs in pmids. They run before the searches. A PMID that does not exist returns a free error row.
๐ Can I filter by publication date?
Yes, with publishedAfter and publishedBefore in YYYY/MM/DD format.
๐ต Do I pay for empty or failed rows?
No. Not-found PMIDs, articles without a title and searches with no results are not charged. If the run hits your maximum charge, it stops cleanly.
๐ข How many articles can one run return?
Up to 1,000,000 on paid plans, but PubMed limits a single query to 9,999 records, so split large topics by date range. Free users get 10 articles per run.
โ๏ธ Does it need an NCBI API key?
No.
๐ซ Why did I get 0 results?
Your query matched nothing on PubMed (try fewer terms or remove the date filter), or you left both searchTerms and pmids empty, which fails the run. Check the run status message.
๐ Want more research and health data? Other PunkRecordsData scrapers
- Clinical Trials Scraper: the trials behind the papers.
- FDA Drug Safety Scraper: recalls, adverse events and labels for the same molecules.
- arXiv Research Papers Scraper: preprints and metadata.
๐ Support
Questions or feature requests: use the Issues tab on this actor page or write to contact.punkrecordsdata@gmail.com. Independent tool, not affiliated with NCBI, NLM or NIH. Only public data. Not medical advice.
Last updated: 2026-10-03