HAL Open Science Scraper
Pricing
from $7.49 / 1,000 result items
HAL Open Science Scraper
Export research papers, theses, and preprints from HAL, France's national open science archive. 3M+ full-text records across every scientific discipline. Filter by domain, author, lab, journal, or year. Pull titles, abstracts, authors, DOIs, PDFs, citations.
Pricing
from $7.49 / 1,000 result items
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
0
Monthly active users
35 minutes ago
Last modified
Categories
Share

๐ HAL Open Science Scraper
๐ Export French open-access research from HAL. 3M+ papers and theses by domain, author, lab, journal.
Export research papers, theses, and preprints from HAL, Frances national open science archive. 3M+ full-text records across every scientific discipline. Filter by domain, author, lab, journal, or year.
Pull titles, abstracts, authors, DOIs, PDFs, citations.
๐ What the HAL Open Science Scraper does
- ๐ฏ Targeted filtering. Use the input schema to narrow results to what you need.
- ๐ฆ Structured output. Clean, typed records with every field documented.
- ๐ Live data. Every run fetches fresh data at runtime, no cached responses.
- ๐ Easy integration. Consume via Apify API, webhooks, or direct dataset export.
- ๐ Scale on demand. Run once or run on a schedule, the same way.
๐ก Why it matters: teams that rely on this source no longer need to babysit a custom crawler. Set up your filters once, get updated data on demand.
๐ Data fields
Each record includes: abstract, abstractEnglish, audience, authors, bookTitle, citationFull, citationRef, conferenceTitle, docId, documentType, doi, domainLabels, domains, edition, halId, halUrl, institutions, isbn, issue, journalIssn, journalTitle, keywords, keywordsEnglish, labs, labStructures, language, openAccess, pages, pdfUrl, peerReviewed, popularLevel, primaryDomain, publicationDate, publisher, scrapedAt, title, titleEnglish, volume, year. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.
โ ๏ธ Good to Know: free users are limited to 10 items per run for preview purposes. Upgrade to Apify paid plans for higher limits.
๐ How to use
- ๐ Create a free account. Sign up at console.apify.com to get $5 in credits.
- ๐ Open the actor. Paste your filters into the input schema in the Apify console.
- โถ๏ธ Click Start. Wait a few seconds for the first records to land.
- ๐ค Export the data. Download JSON/CSV or pipe to webhooks, Google Sheets, or Zapier.
- ๐ Schedule it. Apify Schedules let you rerun on a cron cadence for free.
โฑ๏ธ Total time to first data: about 60 seconds.
๐ Recommended Actors
Pair the HAL Open Science Scraper with related actors:
- ๐ Website Content Crawler - crawl any page at scale
- ๐ Google Search Scraper - harvest SERPs
- ๐ Article Extractor - extract clean article text
- ๐ Google Trends Scraper - capture demand signals
- ๐ธ Screenshot URL - render any page to image
๐ก Pro Tip: browse the complete ParseForge collection for more niche actors.
โ ๏ธ Disclaimer: This actor retrieves data from publicly available sources. You are responsible for complying with the source website's terms of service and applicable laws in your jurisdiction. ParseForge is not affiliated with the data source.
๐ Need Help?
If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.
For faster answers, join our Discord. It's the best place to get support and suggest new actors.