bioRxiv and medRxiv Preprints Scraper avatar

bioRxiv and medRxiv Preprints Scraper

Pricing

from $7.50 / 1,000 results

Go to Apify Store
bioRxiv and medRxiv Preprints Scraper

bioRxiv and medRxiv Preprints Scraper

Track the latest preprints from bioRxiv or medRxiv inside any date window. Returns DOI, title, authors, posting date, category, abstract, version, server, JATS XML link, and license. Useful for literature surveillance, competitive science intelligence, and rapid biomedical research review.

Pricing

from $7.50 / 1,000 results

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

17 hours ago

Last modified

Share

ParseForge Banner

๐Ÿ“‘ bioRxiv Preprints Scraper

๐Ÿš€ Export bioRxiv preprints in seconds. DOIs, titles, authors, dates, categories, abstracts, and JATS XML URLs โ€” direct from the public bioRxiv API.

The bioRxiv Preprints Scraper turns the bioRxiv Public API public endpoint into a clean, structured dataset. It queries the source live, normalizes the response into one row per record, and pushes the result into an Apify dataset you can download or pipe to your warehouse.

Preprints published on bioRxiv (or medRxiv) within the queried date window are covered in a single run, with stable field names and null-safe parsing.

๐ŸŽฏ Target Audience๐Ÿ’ก Primary Use Cases
๐Ÿ”ฌ ResearchersTrack preprints in your field as soon as they post
๐Ÿ’Š PharmaMonitor preprints on a drug or target
๐ŸŽ“ UniversitiesCurate weekly reading lists
๐Ÿค– ML teamsTrain scientific NLP models on the freshest text

๐Ÿ“‹ What the bioRxiv Preprints Scraper does

  • Calls the public bioRxiv Public API endpoint with the parameters you supply.
  • Parses the response and flattens each record into a single dataset row.
  • Casts numeric fields to numbers where applicable for clean spreadsheet imports.
  • Surfaces rate-limit or upstream errors as a single-row error record instead of crashing.
  • Exports to every Apify dataset format supported in the UI.

๐Ÿ’ก Why it matters. The raw bioRxiv Public API response is great for API consumers but awkward for spreadsheets and BI tools. This actor normalizes the shape so the data drops straight into pandas, BigQuery, or a Google Sheet.

๐Ÿ“Š Data fields

Each record includes: abstract, authors, category, date, doi, jatsxml, license, results, scrapedAt, server, title, version. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

โš ๏ธ Good to Know. This actor calls the public bioRxiv Public API endpoint with no authentication required. Upstream rate limits apply; if the source returns a limit notice, you will see it as a single error record in your dataset.

๐Ÿš€ How to use

  1. Click Try for free.
  2. Fill in the input (or leave defaults).
  3. Click Start.
  4. Within seconds, the dataset is ready for download or integration.
ActorWhat it does
ParseForge OurAirports ScraperGlobal airport database.
ParseForge Alpha Vantage ScraperStocks, FX, crypto, and indicators.
ParseForge CurseForge Mods ScraperPublic mod metadata from CurseForge.
ParseForge NBA Stats ScraperPlayer and team stats from NBA.com.

๐Ÿ’ก Pro Tip. Browse the complete ParseForge collection for 900+ production-grade scrapers across business intelligence, real estate, e-commerce, sports, finance, and public records.

Disclaimer. This actor scrapes only publicly available data. ParseForge is not affiliated with, endorsed by, or sponsored by any of the third-party services referenced. Users are responsible for complying with the target site's terms of service and applicable law. Create a free account w/ $5 credit.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.