PubMed Scraper avatar

PubMed Scraper

Pricing

from $1.20 / 1,000 results

Go to Apify Store
PubMed Scraper

PubMed Scraper

Extract biomedical articles from PubMed via the NCBI E-utilities API. 37M+ records: PMID, DOI, title, abstract, authors, journal, MeSH terms, keywords, references. No API key, no browser, no proxy.

Pricing

from $1.20 / 1,000 results

Rating

0.0

(0)

Developer

Aurenic

Aurenic

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

Extract biomedical articles from PubMed via the NCBI E-utilities API. 37M+ records: PMID, DOI, title, abstract, authors, journal, MeSH terms, keywords, references. No API key, no browser, no proxy.

What does PubMed Scraper do?

Scrape PubMed, the U.S. National Library of Medicine's biomedical literature database, in two modes:

  • Search by query — full PubMed syntax: field tags (cancer[Title]), Boolean operators (AND/OR/NOT), quoted phrases, date filters (2024[PDAT]), and publication-type filters. Auto-paginates through results.
  • Fetch by PMID — pass exact PubMed IDs, get full records in batch.

Every result includes the complete article metadata: PMID, DOI, PMC ID, title, journal, volume/issue/pages, publication date, author list with affiliations, full abstract (with labeled sections like "BACKGROUND:", "METHODS:"), MeSH descriptors with qualifiers, keywords, publication types, and reference count.

PubMed is fully public. No API key required — the actor sends the NCBI-required tool and email parameters with every request to stay in good standing.

Output fields

FieldDescription
pmidPubMed ID
doiDOI identifier
pmcIdPubMed Central ID (open-access articles)
titleArticle title
journalFull journal name
journalAbbreviationISO journal abbreviation
volume / issue / pages / issnBibliographic details
publicationDatePublication date
authorsArray of { lastName, firstName, initials, fullName, affiliation }
authorCount / firstAuthor / lastAuthorAuthor summary
abstractFull abstract with section labels
abstractLengthCharacter count
meshTermsArray of { descriptor, qualifiers[] }
meshCountNumber of MeSH descriptors
keywordsAuthor-supplied keywords
publicationTypesReview, Clinical Trial, Meta-Analysis, etc.
referenceCountNumber of cited references
languageArticle language
pubmedUrlDirect PubMed URL

Who is it for?

  • Systematic review teams building PRISMA-compliant article datasets
  • Medical and life-science researchers sourcing literature for meta-analyses
  • AI/ML researchers building biomedical NLP corpora and training sets
  • Competitive intelligence in pharma and biotech tracking publication activity
  • Medical writers and editors finding evidence for regulatory submissions
  • Academic libraries enriching catalogs with abstracts and MeSH data

Pricing

$1.20 per 1,000 results. No subscription.

ResultsCost
100$0.12
1,000$1.20
10,000$12.00

How to use it

  1. Pick a Mode.
  2. For search: enter a Search Query — PubMed syntax supported.
  3. For fetch: enter PMIDs.
  4. Optionally filter by Date Range and Publication Types.
  5. Optionally set your real Contact Email (required by NCBI).
  6. Click Start.

Output example

{
"recordType": "article",
"pmid": "38000000",
"doi": "10.1038/s41586-023-06775-1",
"pmcId": "PMC10705885",
"title": "CRISPR-Cas9 genome editing in human primary T cells",
"journal": "Nature",
"journalAbbreviation": "Nature",
"volume": "624",
"issue": "7992",
"pages": "415-423",
"issn": "0028-0836",
"publicationDate": "2023-12-14",
"authors": [
{ "lastName": "Smith", "firstName": "Jane", "initials": "J", "fullName": "Jane Smith", "affiliation": "Department of Immunology, MIT" },
{ "lastName": "Chen", "firstName": "Wei", "initials": "W", "fullName": "Wei Chen", "affiliation": "Broad Institute" }
],
"authorCount": 12,
"firstAuthor": "Jane Smith",
"lastAuthor": "David Liu",
"abstract": "BACKGROUND: CRISPR-Cas9 has revolutionized genome editing...\n\nMETHODS: We engineered primary human T cells...\n\nRESULTS: Editing efficiency exceeded 95%...",
"abstractLength": 1420,
"meshTerms": [
{ "descriptor": "Gene Editing", "qualifiers": ["methods"] },
{ "descriptor": "CRISPR-Cas Systems", "qualifiers": [] },
{ "descriptor": "T-Lymphocytes", "qualifiers": ["metabolism"] }
],
"meshCount": 14,
"keywords": ["genome editing", "T cells", "CRISPR"],
"publicationTypes": ["Journal Article", "Research Support, N.I.H., Extramural"],
"referenceCount": 48,
"language": "eng",
"pubmedUrl": "https://pubmed.ncbi.nlm.nih.gov/38000000/",
"scrapedAt": "2026-09-25T12:00:00.000Z"
}

Technical details

  • Source: NCBI E-utilities API at https://eutils.ncbi.nlm.nih.gov/entrez/eutils/. Free, no key, no login .
  • Rate limits: 3 req/s anonymous, 10 req/s with a free NCBI API key. The actor defaults to 400ms between requests and backs off on 429 .
  • NCBI policy compliance — every request sends tool= and email= parameters. This is NCBI's documented requirement to avoid IP blocks and to receive early warnings about policy changes.
  • esearch → efetch pipeline — esearch returns PMIDs, efetch returns full XML records in batches of 200.
  • XML parsing via cheerio in xmlMode — extracts structured records from the PubMed XML schema.
  • No browser, no proxy — pure HTTP.

Known limits

  • Rate limit is 3 req/s anonymous. For high-volume runs, register a free NCBI API key to raise it to 10 req/s .
  • NCBI may block datacenter IPs that overload the servers. The actor's tool and email params satisfy NCBI's registration requirement and prevent this — but running at the rate limit is not advised.
  • retmax caps at 10,000 PMIDs per search. For deeper results, paginate with retstart.
  • efetch batches cap at ~200 PMIDs per request. The actor handles this automatically.
  • Full text is not included. PubMed provides abstracts and metadata; full text lives in PubMed Central (PMC) or publisher sites.
  • Retractions and updates are noted in the PubMed record but not surfaced as separate fields.

FAQ

Do I need an API key? No. Anonymous access works at 3 req/s. A free NCBI API key raises the limit to 10 req/s .

Do I need a proxy? No. Datacenter IPs work when sending the required tool and email params.

What query syntax is supported? Full PubMed syntax: field tags ([Title], [Author], [PDAT]), Boolean operators, quoted phrases, and filters. See the PubMed search guide.

How do I search by MeSH term? Use the [MeSH Terms] tag: diabetes[MeSH Terms] AND 2024[PDAT].

How do I search a specific author? Use the [Author] tag: Smith J[Author] AND cancer[Title].

How do I export data? After a run, go to Storage → Export as JSON, CSV, Excel.

Support

Open an issue on the Actor's page for bugs or feature requests.