OpenAlex Scholarly Works Scraper
Pricing
from $6.80 / 1,000 results
OpenAlex Scholarly Works Scraper
Scrape scholarly works with title, DOI, publication year, type, citation count, authors, venue, open access status and a direct link. Search by keyword. Export to JSON, CSV or Excel.
Pricing
from $6.80 / 1,000 results
Rating
0.0
(0)
Developer
Scrapers Lat
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
13 hours ago
Last modified
Categories
Share
OpenAlex Scholarly Works Scraper
Here is one real result, with every field the actor returns (the authorsDetailed array is trimmed here with a note; all values are real, and the AI fields shown come from the optional paid add-ons):
{"id": "https://openalex.org/W1775749144","doi": "10.1016/s0021-9258(19)52451-6","title": "PROTEIN MEASUREMENT WITH THE FOLIN PHENOL REAGENT","publicationYear": 1951,"publicationDate": "1951-11-01","type": "article","citedByCount": 318812,"authors": ["OliverH. Lowry", "NiraJ. Rosebrough", "A. Farr", "RoseJ. Randall"],"venue": "Journal of Biological Chemistry","hostOrganization": "Elsevier BV","isOpenAccess": true,"openAccessStatus": "hybrid","openAccessUrl": "https://www.jbc.org/article/S0021-9258(19)52451-6/pdf","language": "en","url": "https://doi.org/10.1016/s0021-9258(19)52451-6","abstract": "Since 1922 when Wu proposed the use of the Folin phenol reagent for the measurement of proteins (l), a number of modified analytical procedures utilizing this reagent have been reported for the determination of proteins in serum, in antigen-antibody precipitates, and in insulin.","authorsDetailed": [{"name": "OliverH. Lowry","orcid": null,"isCorresponding": false,"institutions": ["Washington University in St. Louis"],"countries": ["US"]}],"concepts": ["Reagent", "Chemistry", "Phenol", "Chromatography", "Organic chemistry"],"topics": ["Glycosylation and Glycoproteins Research", "Muscle metabolism and nutrition", "Cancer and biochemical research"],"keywords": ["Reagent", "Chemistry", "Phenol", "Chromatography", "Organic chemistry"],"fieldsOfStudy": ["Biochemistry, Genetics and Molecular Biology", "Life Sciences"],"referencedWorksCount": 20,"fwci": 48.347,"isRetracted": false,"isParatext": false,"volume": "193","issue": "1","firstPage": "265","lastPage": "275","pdfUrl": "https://www.jbc.org/article/S0021-9258(19)52451-6/pdf","license": "cc-by","funders": ["American Cancer Society"],"sustainableDevelopmentGoals": ["Clean water and sanitation"],"pmid": "14907713","pmcid": null,"source": "OpenAlex","observedAt": "2026-08-14T06:25:00.453Z","aiSummary": "This study discusses the historical use of the Folin phenol reagent for measuring protein levels, highlighting various modified methods developed since its introduction in 1922 for different biological samples.","aiKeywords": ["Protein measurement", "Folin phenol reagent", "Analytical procedures", "Serum proteins", "Insulin determination"],"aiField": "Analytical Chemistry"}
The most complete OpenAlex scholarly works scraper available. It returns every field the OpenAlex work record exposes, including detailed authorship with institutions and countries, open-access status, citation and impact metrics, funders and SDGs, and adds optional AI abstract summaries, keywords and field classification, so you get exactly the research metadata you need.
📥 Input · 📤 Output · 💰 Pricing · ▶️ Examples
Table of contents
- What it does
- Quickstart
- Input reference
- Output reference
- Example output record
- Run via API and CLI
- Fetch results
- Billing and limits
- FAQ and troubleshooting
What it does
The actor searches the OpenAlex catalog of scholarly works by keyword, paginates through the matching works, and writes one normalized record per work to the run's dataset. Each record carries the full OpenAlex metadata (DOI, authors and institutions, venue, open-access status, citation and impact metrics, concepts, topics, funders, SDGs, and identifiers such as PMID and PMCID). Optional paid AI add-ons enrich each work with a plain-English abstract summary, topical keywords, and a field classification.
Coverage is global: OpenAlex indexes hundreds of millions of works across every discipline and country.
Quickstart
Open the actor, paste this into the input, and press Run. It returns 3 works on CRISPR gene editing with all AI add-ons on.
{"searchQuery": "crispr gene editing","maxWorks": 3,"withAiSummary": true,"withAiKeywords": true,"withAiField": true}
Leave the AI add-ons off (default) for raw OpenAlex metadata only. Every input field is optional.
Input reference
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
searchQuery | string | no | machine learning | Keyword to search scholarly works by title and full text, for example machine learning, crispr, climate policy. Leave empty for top works by relevance. |
maxWorks | integer | no | 10 | Maximum number of works to collect. |
withAiSummary | boolean | no | false | Paid add-on. Generates a 1-2 sentence plain-English summary from each work's abstract (skipped when no abstract). Requires a paid Apify plan. Billed only when a summary is produced. |
withAiKeywords | boolean | no | false | Paid add-on. Extracts 5-10 topical keywords per work from its title and abstract. Requires a paid Apify plan. Billed only when keywords are produced. |
withAiField | boolean | no | false | Paid add-on. Classifies each work's academic field or discipline. Requires a paid Apify plan. Billed only when a classification is produced. |
Output reference
One dataset item per work. Types: string, integer, number, boolean, string[], object[], or null when the source value is absent.
| Field | Type | Description |
|---|---|---|
id | string | OpenAlex work ID (URL). |
doi | string | DOI of the work, or null. |
title | string | Work title. |
publicationYear | integer | Year of publication. |
publicationDate | string | Publication date (YYYY-MM-DD). |
type | string | Work type, for example article, book-chapter, dataset. |
citedByCount | integer | Number of times the work has been cited. |
authors | string[] | Author display names in order. |
venue | string | Source/journal name. |
hostOrganization | string | Publisher or host organization. |
isOpenAccess | boolean | true when the work is open access. |
openAccessStatus | string | Open-access status, for example gold, hybrid, green, closed. |
openAccessUrl | string | URL to the open-access copy, or null. |
language | string | Language code, for example en. |
url | string | Canonical link to the work (DOI when available). |
abstract | string | Reconstructed abstract text, or null. |
authorsDetailed | object[] | Per-author detail: name, orcid, isCorresponding, institutions, countries. |
concepts | string[] | OpenAlex concept labels. |
topics | string[] | OpenAlex topic labels. |
keywords | string[] | OpenAlex keyword labels. |
fieldsOfStudy | string[] | Broad fields of study. |
referencedWorksCount | integer | Number of works referenced. |
fwci | number | Field-Weighted Citation Impact, or null. |
isRetracted | boolean | true when the work is retracted. |
isParatext | boolean | true when the work is paratext (front matter, and similar). |
volume | string | Volume, or null. |
issue | string | Issue, or null. |
firstPage | string | First page, or null. |
lastPage | string | Last page, or null. |
pdfUrl | string | Direct PDF URL when available, else null. |
license | string | License, for example cc-by, or null. |
funders | string[] | Funding organizations. |
sustainableDevelopmentGoals | string[] | UN SDG labels linked to the work. |
pmid | string | PubMed ID when available, else null. |
pmcid | string | PubMed Central ID when available, else null. |
source | string | Data source name. Always OpenAlex. |
observedAt | string | ISO 8601 timestamp of when the record was collected. |
aiSummary | string | AI abstract summary. Present only when withAiSummary is on and an abstract exists. |
aiKeywords | string[] | AI-extracted keywords. Present only when withAiKeywords is on. |
aiField | string | AI field classification. Present only when withAiField is on. |
On a failed run, a single item with a populated error field (plus source and observedAt) is written instead.
Example output record
Real record from a live run (input {"searchQuery":"crispr gene editing","maxWorks":3}, AI add-ons off). Some fields omitted here for brevity; all values are real:
{"id": "https://openalex.org/W1775749144","doi": "10.1016/s0021-9258(19)52451-6","title": "PROTEIN MEASUREMENT WITH THE FOLIN PHENOL REAGENT","publicationYear": 1951,"type": "article","citedByCount": 318812,"authors": ["OliverH. Lowry", "NiraJ. Rosebrough", "A. Farr", "RoseJ. Randall"],"venue": "Journal of Biological Chemistry","hostOrganization": "Elsevier BV","isOpenAccess": true,"openAccessStatus": "hybrid","language": "en","url": "https://doi.org/10.1016/s0021-9258(19)52451-6","fieldsOfStudy": ["Biochemistry, Genetics and Molecular Biology", "Life Sciences"],"fwci": 48.347,"license": "cc-by","funders": ["American Cancer Society"],"pmid": "14907713","source": "OpenAlex","observedAt": "2026-08-14T06:25:00.453Z"}
Run via API and CLI
Start a run and wait for it to finish, then read the dataset. Replace <TOKEN> with your Apify API token.
Run synchronously and get dataset items in one call:
curl -X POST "https://api.apify.com/v2/acts/scrapers_lat~openalex-scraper/run-sync-get-dataset-items?token=<TOKEN>" \-H "Content-Type: application/json" \-d '{"searchQuery":"machine learning","maxWorks":25}'
Start a run asynchronously with AI add-ons:
curl -X POST "https://api.apify.com/v2/acts/scrapers_lat~openalex-scraper/runs?token=<TOKEN>" \-H "Content-Type: application/json" \-d '{"searchQuery":"climate policy","maxWorks":50,"withAiSummary":true,"withAiField":true}'
Apify CLI:
apify call scrapers_lat/openalex-scraper \--input '{"searchQuery":"crispr"}'
Fetch results
Every run writes to a dataset. Fetch items as JSON, CSV, or Excel by changing format:
# JSONcurl "https://api.apify.com/v2/datasets/<DATASET_ID>/items?token=<TOKEN>&clean=true&format=json"# CSVcurl "https://api.apify.com/v2/datasets/<DATASET_ID>/items?token=<TOKEN>&clean=true&format=csv"# Paginate large datasetscurl "https://api.apify.com/v2/datasets/<DATASET_ID>/items?token=<TOKEN>&offset=1000&limit=1000"
<DATASET_ID> is returned as defaultDatasetId in the run object. Use offset and limit to page through large result sets. clean=true drops empty and internal fields.
Billing and limits
- Pay per result. You are charged per work returned (
resultevent). See the pricing tab for the current per-result price. - AI add-ons are billed separately.
ai_summary,ai_keywords, andai_fieldare each charged per work only when they produce usable output, and only on paid Apify plans. - No charge on failure. If a request fails, the actor writes a single item with a populated
errorfield and does not charge for it. Empty runs cost nothing. - Spend cap respected. Set
maxTotalChargeUsdon the run; once reached, the actor stops emitting and charging further billable results. - Free Apify plans are capped at 10 records per run and cannot use the AI add-ons. Upgrade for higher
maxWorksand AI enrichment.
FAQ and troubleshooting
A run returned 0 records. Why? The keyword matched no works in OpenAlex. Try a broader query. Zero-result runs are not charged.
Do I need an API key? No. OpenAlex is an open catalog and the actor reads it directly. No account or key is required.
What do the AI add-ons cost? Each add-on is billed per work only when it returns usable output, and requires a paid Apify plan. Leave them off for raw OpenAlex metadata.
Why is abstract null?
OpenAlex does not hold a reconstructable abstract for every work. Missing source values are returned as null, never invented. The AI summary is skipped when there is no abstract.
What is fwci?
Field-Weighted Citation Impact, a normalized measure of how a work's citations compare to the field average. It is null when OpenAlex has not computed it.
Is this an official OpenAlex tool? No. This actor is independent and has no affiliation with OpenAlex or OurResearch. It reads the open OpenAlex catalog. Use it in accordance with the OpenAlex terms of use.
Related scrapers
- Crossref Works Scraper: Scholarly metadata and DOIs from Crossref.
- arXiv Papers Scraper: arXiv preprints with authors and abstracts.
- ClinicalTrials.gov Scraper: Clinical trial records worldwide.
- GBIF Species Occurrence Scraper: Global biodiversity species records.
- Hugging Face Models Scraper: Machine-learning model metadata.
- Google News Scraper: News articles by keyword.
More scrapers at scrapers.lat
Built and maintained by scrapers.lat, where we publish scrapers for US and Latin American public platforms: company registries, government data, finance, e-commerce and more. Browse the catalog or request a custom scraper at scrapers.lat.
Independent tool, not affiliated with OpenAlex or OurResearch. Reads the open OpenAlex catalog. Use in accordance with the OpenAlex terms of use.
