PubMed Articles Scraper avatar

PubMed Articles Scraper

Pricing

from $1.33 / 1,000 results

Go to Apify Store
PubMed Articles Scraper

PubMed Articles Scraper

Search PubMed and get structured article metadata: title, authors, journal, publication dates, volume and pages, DOI, PMC ID, publication types and citation counts. Keyless NCBI E-utilities API.

Pricing

from $1.33 / 1,000 results

Rating

0.0

(0)

Developer

Farhan Febrian Nauval

Farhan Febrian Nauval

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

PubMed Articles Scraper — Search Results with Full Citation Metadata

Search PubMed and get structured article records: title, full author list, journal and ISSNs, publication dates, volume/issue/pages, DOI and PMC ID, publication types and citation counts.

Keyless NCBI E-utilities API. No login, no browser, no proxy needed.

Input

FieldTypeDefaultDescription
searchQueriesarrayrequiredFull PubMed syntax, e.g. crispr AND cancer[Title]
publishedFrom / publishedTostring—YYYY/MM/DD
dateTypestringpdatWhich date the filter matches — see below
sortstringrelevancepub_date, Author, JournalName
maxItemsPerQueryinteger200Capped by the API at 9,999
apiKeystring (secret)—Free from NCBI; raises 3 req/s to 10
contactEmailstring—NCBI asks callers to identify themselves

Output

{
"_input": "crispr AND cancer",
"_source": "S1-eutils-esummary",
"_scrapedAt": "2026-09-09T13:44:12Z",
"pmid": "42711491",
"url": "https://pubmed.ncbi.nlm.nih.gov/42711491/",
"title": "Chemogenomic maps reveal a PRDX1-dependent iron-damage axis…",
"authors": ["O'Loughlin TA", "Arab A", …],
"authorCount": 14,
"firstAuthor": "O'Loughlin TA", "lastAuthor": "Bhatt DM",
"journal": "Nat Chem Biol",
"journalFull": "Nature chemical biology",
"issn": "1552-4450", "eIssn": "1552-4469", "nlmUniqueId": "101231976",
"pubDate": "2026 Sep 8", // ← what a date filter matches
"ePubDate": "2026 Sep 8",
"sortPubDate": "2026/09/08 00:00", // ← PubMed's sort key; see below
"docDate": null,
"volume": "", "issue": "", "pages": "",
"language": ["eng"],
"publicationTypes": ["Journal Article"],
"doi": "10.1038/s41589-026-02312-z",
"doiUrl": "https://doi.org/10.1038/s41589-026-02312-z",
"pmcId": null,
"elocationId": "doi: 10.1038/s41589-026-02312-z",
"articleIds": { "pubmed": "42711491", "doi": "…" },
"pmcRefCount": 12,
"recordStatus": "PubMed - as supplied by publisher",
"publicationStatus": "aheadofprint"
}

Three things worth knowing

pubDate and sortPubDate are different dates, and the filter matches the first. PubMed's sort key reflects later revisions: a 2010 article updated in 2024 sorts as 2024. Filtering a run to 2010 returned rows whose pubDate was 2010 for 37 of 40 — while their sortPubDate was mostly 2013 and 2014. Compare against pubDate, or the filter will look broken when it is working exactly as specified. dateType lets you filter on the Entrez (edat) or modification (mdat) date instead.

PubMed serves at most 9,999 records per query. retstart cannot exceed 9998, and retmax is silently clamped: ask for 10,000 and you get 9,999, no warning. Narrowing by date range is the way past it — the actor logs the true match count and says so when you have asked for more than can be served.

The error path is not valid JSON. Overrunning the ceiling returns HTTP 200 with an ERROR key, and the message embeds a raw newline inside a JSON string. Strict json.loads refuses it outright with Invalid control character, so the actual explanation is lost and the failure looks like a transport error. The success path parses fine — meaning a parser can pass every test and still fail on the first real failure. This actor parses leniently and reports the message.

Rate limits and courtesy

NCBI documents 3 requests a second without a key and 10 with one, and asks callers to identify themselves so they can get in touch about unusual usage rather than simply blocking it. The actor paces itself to the documented rate, sends a tool identifier, and passes your contactEmail when you give one. Neither is authentication; both are the terms of use.

Errors

_errorMeaning
invalid_inputEmpty query
no_resultsThe search ran and matched nothing
api_errorA 200 carrying an ERROR field — usually the 9,999 ceiling
unexpected_shapeA 200 without esearchresult or result
blockedEvery TLS profile was refused
network_errorThe ladder never reached the server

If every query fails, the run itself fails rather than reporting success over an empty dataset.