HAL Open Science Scraper avatar

HAL Open Science Scraper

Pricing

from $7.49 / 1,000 result items

Go to Apify Store
HAL Open Science Scraper

HAL Open Science Scraper

Export research papers, theses, and preprints from HAL, France's national open science archive. 3M+ full-text records across every scientific discipline. Filter by domain, author, lab, journal, or year. Pull titles, abstracts, authors, DOIs, PDFs, citations.

Pricing

from $7.49 / 1,000 result items

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

0

Monthly active users

35 minutes ago

Last modified

Share

ParseForge Banner

๐Ÿš€ HAL Open Science Scraper

๐Ÿš€ Export French open-access research from HAL. 3M+ papers and theses by domain, author, lab, journal.

Export research papers, theses, and preprints from HAL, Frances national open science archive. 3M+ full-text records across every scientific discipline. Filter by domain, author, lab, journal, or year.

Pull titles, abstracts, authors, DOIs, PDFs, citations.

๐Ÿ“‹ What the HAL Open Science Scraper does

  • ๐ŸŽฏ Targeted filtering. Use the input schema to narrow results to what you need.
  • ๐Ÿ“ฆ Structured output. Clean, typed records with every field documented.
  • ๐Ÿ”„ Live data. Every run fetches fresh data at runtime, no cached responses.
  • ๐Ÿ”Œ Easy integration. Consume via Apify API, webhooks, or direct dataset export.
  • ๐Ÿ“Š Scale on demand. Run once or run on a schedule, the same way.

๐Ÿ’ก Why it matters: teams that rely on this source no longer need to babysit a custom crawler. Set up your filters once, get updated data on demand.

๐Ÿ“Š Data fields

Each record includes: abstract, abstractEnglish, audience, authors, bookTitle, citationFull, citationRef, conferenceTitle, docId, documentType, doi, domainLabels, domains, edition, halId, halUrl, institutions, isbn, issue, journalIssn, journalTitle, keywords, keywordsEnglish, labs, labStructures, language, openAccess, pages, pdfUrl, peerReviewed, popularLevel, primaryDomain, publicationDate, publisher, scrapedAt, title, titleEnglish, volume, year. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

โš ๏ธ Good to Know: free users are limited to 10 items per run for preview purposes. Upgrade to Apify paid plans for higher limits.

๐Ÿš€ How to use

  1. ๐Ÿ“ Create a free account. Sign up at console.apify.com to get $5 in credits.
  2. ๐Ÿ” Open the actor. Paste your filters into the input schema in the Apify console.
  3. โ–ถ๏ธ Click Start. Wait a few seconds for the first records to land.
  4. ๐Ÿ“ค Export the data. Download JSON/CSV or pipe to webhooks, Google Sheets, or Zapier.
  5. ๐Ÿ”„ Schedule it. Apify Schedules let you rerun on a cron cadence for free.

โฑ๏ธ Total time to first data: about 60 seconds.

Pair the HAL Open Science Scraper with related actors:

๐Ÿ’ก Pro Tip: browse the complete ParseForge collection for more niche actors.

โš ๏ธ Disclaimer: This actor retrieves data from publicly available sources. You are responsible for complying with the source website's terms of service and applicable laws in your jurisdiction. ParseForge is not affiliated with the data source.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.