Europe PMC Literature Scraper avatar

Europe PMC Literature Scraper

Pricing

from $27.60 / 1,000 results

Go to Apify Store
Europe PMC Literature Scraper

Europe PMC Literature Scraper

Scrape Europe PMC for biomedical research papers. Search by title, author, MeSH terms, journal. Get DOI, abstract, full-text URLs, citations, references, open-access status. No API key required.

Pricing

from $27.60 / 1,000 results

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

4

Total users

0

Monthly active users

21 hours ago

Last modified

Share

ParseForge Banner

๐Ÿงฌ Europe PMC Literature Scraper

๐Ÿš€ Export the biomedical literature index in seconds. Search 40+ million records across PubMed, PubMed Central, life-science preprints, agricultural literature, and patents. Filter by title, author, MeSH term, DOI, journal, open access, or free-text. No API key, no registration.

The Europe PMC Literature Scraper wraps the official Europe PMC REST API (ebi.ac.uk/europepmc/webservices/rest/search) and returns one row per article with 40+ fields, including DOI, PMID, PMCID, abstract, full-text URLs, MeSH terms, keywords, journal, citation count, open-access status, and licensing. The underlying corpus is published by Europe PMC, the European mirror of PubMed Central, maintained by EMBL-EBI and funded by 32 life-science research funders worldwide.

The index covers MEDLINE/PubMed, PubMed Central (full text), Agricola (USDA agricultural literature), bioRxiv and medRxiv preprints, CTX patents, and Europe PMC-curated content. Free-text and field-qualified queries (TITLE, AUTH, MESH, DOI, PMID, AFFILIATION, JOURNAL) compose freely with boolean operators. This Actor returns structured records ready to download as CSV, Excel, JSON, or XML.

๐ŸŽฏ Target Audience๐Ÿ’ก Primary Use Cases
Biomedical researchers, systematic-review teams, bibliometrics analysts, pharma intelligence, scientific publishers, science journalists, OA advocacy, ML training pipelinesLiterature reviews, MeSH-term mining, author publication tracking, journal impact studies, drug-target evidence harvesting, training-set assembly

๐Ÿ“‹ What the Europe PMC Scraper does

One programmable interface to the full Europe PMC search service:

  • ๐Ÿ” Field-qualified queries. TITLE:, AUTH:, AFFILIATION:, JOURNAL:, MESH:, DOI:, PMID:, PMCID:, OPEN_ACCESS:, plus boolean operators (AND, OR, NOT) and quoted phrases.
  • ๐Ÿ“š Three response shapes. core returns the full record with abstract, full-text URLs, and metadata. lite returns compact fields. idlist returns IDs only for ultra-fast scans.
  • โฑ๏ธ Sort options. Relevance (default), newest first, oldest first, or most cited.
  • ๐Ÿ” Cursor-mark pagination. Fully automatic. Walks the entire result set efficiently for large queries.

Output captures the publication metadata (PMID, PMCID, DOI, source, journal title, ISSN, volume, issue, page info, publication year and date), full author list, abstract text, affiliation, language, publication types, MeSH headings, keywords, grant count, citation count, full-text URLs, license, open-access flag, and indexing dates.

๐Ÿ’ก Why it matters: Europe PMC is the deepest open-access biomedical literature index in the world. The web UI is great for one-off lookups, but systematic reviews, bibliometric studies, and ML training-set assembly need flat rows. This Actor turns the search service into a downloadable dataset in one run.

๐Ÿ“Š Data fields

Each record includes: abstractText, affiliation, authorList, authorString, citedByCount, dateOfRevision, doi, firstIndexDate, firstPublicationDate, fullTextUrls, grantsCount, hasBook, hasDbCrossReferences, hasPDF, hasReferences, hasSuppl, hasTextMinedTerms, id, inEPMC, inPMC, isOpenAccess, issue, journalIssn, journalTitle, journalVolume, keywords, language, license, meshTerms, pageInfo, pmcid, pmid, pubDate, pubYear, publicationStatus, publicationTypes, scrapedAt, source, title, url. All 40 field names come from a real production run, so what you see here is what lands in your dataset.

๐Ÿš€ How to use

  1. ๐Ÿ“ Sign up. Create a free account with $5 credit (takes 2 minutes).
  2. ๐ŸŒ Open the Actor. Go to the Europe PMC Literature Scraper page on the Apify Store.
  3. ๐Ÿ” Build a query. Free-text or use field qualifiers (AUTH:"Doudna J" AND CRISPR).
  4. ๐Ÿ“š Pick a response shape. core for full metadata, lite for compact, idlist for IDs only.
  5. ๐Ÿš€ Run it. Click Start and let the Actor collect your data.
  6. ๐Ÿ“ฅ Download. Grab your results in the Dataset tab as CSV, Excel, JSON, or XML.

โฑ๏ธ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.

๐Ÿ’ก Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

โš ๏ธ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by Europe PMC, EMBL-EBI, the European Bioinformatics Institute, the National Center for Biotechnology Information, or any of the 32 funders supporting Europe PMC. All trademarks mentioned are the property of their respective owners. Only publicly available open data from the official Europe PMC REST API is collected.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.