OpenAlex Scholarly Works Scraper avatar

OpenAlex Scholarly Works Scraper

Pricing

Pay per event

Go to Apify Store
OpenAlex Scholarly Works Scraper

OpenAlex Scholarly Works Scraper

Export academic works, authors, institutions, sources, and concepts from OpenAlexs open catalog of 250M+ scholarly records. Successor to Microsoft Academic Graph. Filter by author, concept, year, open access status, or affiliation.

Pricing

Pay per event

Rating

5.0

(1)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

16

Total users

2

Monthly active users

16 hours ago

Last modified

Share

ParseForge Banner

🎓 OpenAlex Scholarly Works Scraper

🚀 Export academic works, authors, institutions, and more from OpenAlex in seconds. Filter by search query, entity type, or custom filters. No coding, no API keys required.

The OpenAlex Scholarly Works Scraper connects to OpenAlex, the free and open catalog of 250M+ scholarly records that succeeded Microsoft Academic Graph. It supports 7 entity types: works, authors, institutions, sources, concepts, publishers, and funders. Each record includes 30+ structured fields with titles, DOIs, citation counts, open access status, author details, institutional affiliations, and more. Whether you need 10 papers for a quick lookup or millions of records for a large-scale bibliometric study, this tool handles it efficiently.

Built for researchers conducting literature reviews, bibliometricians analyzing citation networks, university administrators tracking institutional output, and data teams building scholarly knowledge graphs. The scraper uses the OpenAlex API with support for free-text search and the full OpenAlex filter syntax. Providing a contact email puts your requests in the "polite pool" for faster processing.

Target AudienceUse Cases
Academic ResearchersLiterature reviews, citation analysis
BibliometriciansCitation network mapping, impact studies
University AdministratorsInstitutional output tracking
Data ScientistsKnowledge graph construction, NLP corpus building
Funding AgenciesResearch output assessment, grant evaluation
Library ScientistsCollection development, trend analysis

📋 What the OpenAlex Scholarly Works Scraper does

  • 📝 Extracts scholarly work metadata including titles, abstracts, DOIs, publication dates, and citation counts for bibliometric analysis
  • 👥 Collects author profiles with names, ORCID IDs, institutional affiliations, and publication histories
  • 🏫 Gathers institution data including names, types, locations, and research output statistics
  • 📰 Pulls source information for journals, conferences, and repositories with ISSN, publisher, and open access details
  • 🔗 Captures concept and topic data for subject classification and research trend analysis
  • 📊 Tracks open access status with OA type, OA URL, and license information for each work

The scraper queries the OpenAlex API with your search terms and optional filters, handles cursor-based pagination, and processes results efficiently. The OpenAlex filter syntax supports field-level filtering like publication_year:2024,is_oa:true,authorships.institutions.country_code:US for precise targeting.

💡 Why it matters: OpenAlex is the largest free scholarly database, covering 250M+ works, 90M+ authors, and 100K+ institutions. This scraper gives you structured access to this data without writing API integration code.

📊 Data fields

Each record includes: affiliations, citedByCount, displayName, entity, hIndex, i10Index, lastKnownInstitution, openalexId, orcid, scrapedAt, title, twoYearMeanCitedness, url, worksApiUrl, worksCount. All 15 field names come from a real production run, so what you see here is what lands in your dataset.

⚠️ Good to Know: Providing your email address puts your requests in OpenAlex's "polite pool" for faster rate limits. The filter syntax supports dozens of fields. Free users are automatically limited to 10 items per run.

🚀 How to use

  1. Create a free Apify account - Sign up here (includes free credits)
  2. Open the OpenAlex Scholarly Works Scraper - Navigate to the Actor page and click "Start"
  3. Choose your entity type - Select works, authors, institutions, or another entity type
  4. Set your search and filters - Enter a search query and optional OpenAlex filters
  5. Run and download - Click "Start", wait for completion, then export as JSON, CSV, or Excel

⏱️ First results appear in under 10 seconds. A typical run of 100 records completes in about 30 seconds.

ActorDescription
📚 PubMed Citation ScraperExtract citation data and metadata from PubMed biomedical literature
📖 PLOS Journals ScraperCollect article data from PLOS ONE and other PLOS journals
🧬 Crossref ScraperCollect DOI metadata and citation information from Crossref
📰 medRxiv ScraperExtract health sciences preprint data from medRxiv
📄 Semantic Scholar ScraperQuery the Semantic Scholar API for academic paper data

💡 Pro Tip: Use OpenAlex to find papers by topic, then cross-reference with the Crossref Scraper for detailed citation metadata and reference lists.

Disclaimer: This Actor is provided as-is, without warranty. It is not affiliated with or endorsed by OpenAlex or OurResearch. Use it responsibly and in compliance with applicable terms of service. The authors are not responsible for how the collected data is used. Always verify data accuracy for critical applications.

🆘 Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.