medRxiv Scraper avatar

medRxiv Scraper

Pricing

Pay per event

Go to Apify Store
medRxiv Scraper

medRxiv Scraper

Extract comprehensive preprint data from medRxiv, including titles, authors, abstracts, full text, DOIs, citations, and metadata. Automate access to health-science preprints with structured outputs, ideal for researchers and analysts who need reliable, large-scale article data without manual work.

Pricing

Pay per event

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

10

Total users

1

Monthly active users

15 hours ago

Last modified

Share

ParseForge Banner

๐Ÿงฌ medRxiv Preprint Scraper

๐Ÿš€ Scrape medRxiv health-science preprints in seconds. Filter by topic, subject collection, date range, or author. No API key, no registration, no manual CSV wrangling.

medRxiv is the leading open preprint server for the health sciences, hosting clinical research, epidemiology, public health, and biomedical work months before formal peer review. This Actor turns any medRxiv search (keyword, subject collection, posted-date window, or author) into a structured dataset of full preprint records, complete with title, all authors, posting date, DOI, abstract, full text, PDF link, license, funding statement, competing-interest declarations, data-availability statement, and any data or code repository link the authors disclose. The output drops straight into Google Sheets, BigQuery, Postgres, Notion, or any other tool your team already uses.

Preprint data is hard to harvest at scale. medRxiv exposes a faceted search interface but no public bulk API for researchers; the site sits behind Cloudflare and a Varnish edge that rate-limits direct datacenter traffic. This Actor closes the gap: pick a search query, narrow it with subject collection, date range, or author filters, set how many records you want, and the data lands in your dataset within minutes. Systematic reviewers, clinical research teams, public-health analysts, bibliometric researchers, science journalists, and AI training-set curators all use this kind of feed for living reviews, evidence surveillance, citation tracking, and topic modelling on emerging health research.

๐Ÿ‘ฅ Target audience๐ŸŽฏ Primary use case
Systematic reviewers and meta-analystsRun living searches across health-science preprints
Clinical research and trial teamsTrack new evidence in a therapeutic area in near real time
Public-health analysts and epidemiologistsSurveil outbreak, vaccine, and intervention literature as it lands
Bibliometric and science-policy researchersBuild datasets on topic emergence, author networks, and funding patterns
Science journalists and editorsSpot newsworthy preprints in a chosen field within hours of posting
AI/ML and NLP teamsCurate domain-specific corpora of biomedical text for training and evaluation

๐Ÿ“‹ What the medRxiv Preprint Scraper does

  • ๐Ÿ” Any keyword search. Drive results from a free-text query just like the medRxiv search box.
  • ๐Ÿ—‚๏ธ 50 subject collections. Restrict to a specific medical specialty (Epidemiology, Cardiovascular Medicine, Infectious Diseases, HIV/AIDS, Public and Global Health, and 45 more).
  • ๐Ÿ“… Date range filter. Slice by posting date with dateFrom and dateTo (YYYY-MM-DD).
  • ๐Ÿ‘ค Author filter. Narrow to preprints where a given author name appears.
  • ๐Ÿ“ฐ Full record per preprint. Title, full author list, posting date, DOI, abstract, full text, PDF URL, license, funding, competing interests, data-availability statement, and disclosed data or code URLs.
  • ๐Ÿ” Sort order. Choose newest first, oldest first, or relevance-ranked.

Each record returns the article URL, title, author list (with affiliation metadata where exposed), posting date, DOI, abstract, full text body, subject area assignment from medRxiv, PDF link, citation string, license text, funding statement, competing-interest statement, author declarations block, data-availability statement, disclosed data or code repository URL, supplementary materials, and the scrape timestamp.

๐Ÿ’ก Why it matters: preprints often surface findings weeks or months before journal publication. In fast-moving health topics (pandemics, vaccines, drug repurposing) waiting for the indexed-in-MEDLINE version means working from stale evidence. This Actor lets analysts track the live preprint frontier without manual click-through.

๐Ÿ“Š Data fields

Each record includes: abstract, authorDeclarations, authorDetails, authors, citationInformation, competingInterestStatement, dataAvailability, dataCodeUrl, doi, fullText, fundingStatement, licenseInformation, pdfUrl, publicationDate, scrapedTimestamp, subjectAreas, title, url. All 18 field names come from a real production run, so what you see here is what lands in your dataset.

โš ๏ธ Good to Know: You can use either startUrl or searchQuery, but not both at the same time. If you provide a startUrl, the searchQuery and orderBy fields are ignored. Free users are automatically limited to 10 items per run.

๐Ÿš€ How to use

  1. ๐Ÿ” Sign up. Create a free Apify account (no credit card needed for the free tier).
  2. ๐Ÿ”Ž Open the Actor. Search the Apify Store for "medRxiv" or open this page.
  3. ๐Ÿงฉ Fill in input. Either paste a medRxiv search URL into startUrl, or set searchQuery, subjectCollection, dateFrom / dateTo, and author as needed.
  4. โ–ถ๏ธ Click Start. The Actor uses Apify Residential US proxies to load and parse search and detail pages.
  5. ๐Ÿ“ฅ Export. Download the dataset as JSON, CSV, or Excel, or push directly to Google Sheets, BigQuery, Webhook, or S3.

โฑ๏ธ Total time: about 2 minutes from sign-up to first export.

๐Ÿ’ก Pro Tip: browse the complete ParseForge collection for more research, scholarly, and health-data scrapers.

โš ๏ธ Disclaimer: This Actor extracts publicly accessible preprint data from medRxiv.org for legitimate research, evidence-surveillance, and analytics purposes. It does not collect personal user data beyond what authors and institutions publicly publish on their preprints. Use of this Actor is at your own risk and subject to medRxiv.org's terms of service and the per-preprint Creative Commons or CC0 license. ParseForge is not affiliated with, endorsed by, or sponsored by medRxiv, Cold Spring Harbor Laboratory, BMJ, or Yale University.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.