PubMed Articles Scraper
Pricing
from $1.33 / 1,000 results
PubMed Articles Scraper
Search PubMed and get structured article metadata: title, authors, journal, publication dates, volume and pages, DOI, PMC ID, publication types and citation counts. Keyless NCBI E-utilities API.
Pricing
from $1.33 / 1,000 results
Rating
0.0
(0)
Developer
Farhan Febrian Nauval
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
PubMed Articles Scraper — Search Results with Full Citation Metadata
Search PubMed and get structured article records: title, full author list, journal and ISSNs, publication dates, volume/issue/pages, DOI and PMC ID, publication types and citation counts.
Keyless NCBI E-utilities API. No login, no browser, no proxy needed.
Input
| Field | Type | Default | Description |
|---|---|---|---|
searchQueries | array | required | Full PubMed syntax, e.g. crispr AND cancer[Title] |
publishedFrom / publishedTo | string | — | YYYY/MM/DD |
dateType | string | pdat | Which date the filter matches — see below |
sort | string | relevance | pub_date, Author, JournalName |
maxItemsPerQuery | integer | 200 | Capped by the API at 9,999 |
apiKey | string (secret) | — | Free from NCBI; raises 3 req/s to 10 |
contactEmail | string | — | NCBI asks callers to identify themselves |
Output
{"_input": "crispr AND cancer","_source": "S1-eutils-esummary","_scrapedAt": "2026-09-09T13:44:12Z","pmid": "42711491","url": "https://pubmed.ncbi.nlm.nih.gov/42711491/","title": "Chemogenomic maps reveal a PRDX1-dependent iron-damage axis…","authors": ["O'Loughlin TA", "Arab A", …],"authorCount": 14,"firstAuthor": "O'Loughlin TA", "lastAuthor": "Bhatt DM","journal": "Nat Chem Biol","journalFull": "Nature chemical biology","issn": "1552-4450", "eIssn": "1552-4469", "nlmUniqueId": "101231976","pubDate": "2026 Sep 8", // ← what a date filter matches"ePubDate": "2026 Sep 8","sortPubDate": "2026/09/08 00:00", // ← PubMed's sort key; see below"docDate": null,"volume": "", "issue": "", "pages": "","language": ["eng"],"publicationTypes": ["Journal Article"],"doi": "10.1038/s41589-026-02312-z","doiUrl": "https://doi.org/10.1038/s41589-026-02312-z","pmcId": null,"elocationId": "doi: 10.1038/s41589-026-02312-z","articleIds": { "pubmed": "42711491", "doi": "…" },"pmcRefCount": 12,"recordStatus": "PubMed - as supplied by publisher","publicationStatus": "aheadofprint"}
Three things worth knowing
pubDate and sortPubDate are different dates, and the filter matches the first. PubMed's sort
key reflects later revisions: a 2010 article updated in 2024 sorts as 2024. Filtering a run to 2010
returned rows whose pubDate was 2010 for 37 of 40 — while their sortPubDate was mostly 2013 and
2014. Compare against pubDate, or the filter will look broken when it is working exactly as
specified. dateType lets you filter on the Entrez (edat) or modification (mdat) date instead.
PubMed serves at most 9,999 records per query. retstart cannot exceed 9998, and retmax is
silently clamped: ask for 10,000 and you get 9,999, no warning. Narrowing by date range is the way
past it — the actor logs the true match count and says so when you have asked for more than can be
served.
The error path is not valid JSON. Overrunning the ceiling returns HTTP 200 with an ERROR key,
and the message embeds a raw newline inside a JSON string. Strict json.loads refuses it outright
with Invalid control character, so the actual explanation is lost and the failure looks like a
transport error. The success path parses fine — meaning a parser can pass every test and still fail
on the first real failure. This actor parses leniently and reports the message.
Rate limits and courtesy
NCBI documents 3 requests a second without a key and 10 with one, and asks callers to identify
themselves so they can get in touch about unusual usage rather than simply blocking it. The actor
paces itself to the documented rate, sends a tool identifier, and passes your contactEmail when
you give one. Neither is authentication; both are the terms of use.
Errors
_error | Meaning |
|---|---|
invalid_input | Empty query |
no_results | The search ran and matched nothing |
api_error | A 200 carrying an ERROR field — usually the 9,999 ceiling |
unexpected_shape | A 200 without esearchresult or result |
blocked | Every TLS profile was refused |
network_error | The ladder never reached the server |
If every query fails, the run itself fails rather than reporting success over an empty dataset.