Crossref Works Scraper — DOIs, Authors & Citations
Pricing
Pay per event
Crossref Works Scraper — DOIs, Authors & Citations
Search Crossref's 150M+ DOI registry by bibliographic query and metadata filters, and export title, publisher, journal, full author list with ORCID and affiliation, citation counts, funder/award data, license, ISSN, and abstract as clean JSON, CSV, or Excel.
Pricing
Pay per event
Rating
0.0
(0)
Developer
DevilScrapes
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 hours ago
Last modified
Categories
Share
🎯 What this scrapes
Crossref is the DOI registration agency behind most scholarly publishing — 150M+ journal articles, book chapters, datasets, preprints, and proceedings, each with publisher-submitted metadata exposed through one keyless JSON API. The catch is cursor-only deep paging plus metadata that nests authors, funders, and licenses several levels down. This Actor walks Crossref's cursor pages for you and flattens every work into one row — authors keep their ORCID and affiliation, funders keep their award numbers — so a scholarly-output tracker or funder-compliance report lands straight in a spreadsheet.
🔥 What we handle for you
walks Crossref's cursor paging so there is no ~10 000-record offset ceiling
flattens nested author/funder/license blocks into clean rows — authors keep ORCID and affiliation, funders keep award numbers
retries transient 429/5xx responses with backoff instead of failing the whole run
💡 Use cases
- Track a competitor lab's or institution's new publications as they're registered with Crossref.
- Pull citation counts for a grant or tenure report without hand-rolling the API.
- Check funder compliance — confirm published works credit the right funder and award number.
- Build a publisher- or journal-level dataset of works, ISSNs, and licenses for a research-ops pipeline.
⚙️ How to use it
- Click Try for free at the top of the page.
- Fill in the input form — most fields have sensible defaults.
- Click Start. Output streams into the run's dataset.
- Export from Storage → Dataset as JSON, CSV, or Excel — or fetch via the API.
📥 Input
| Field | Type | Required | Default | Notes |
|---|---|---|---|---|
query | string | no | 'climate change adaptation' | Free-text bibliographic search (title, author, journal, year all matched together). Leave empty if you are filtering… |
fromPublicationDate | string | no | '2020-01-01' | Only works published on or after this date, as YYYY-MM-DD. Leave empty for no floor. |
untilPublicationDate | string | no | '—' | Only works published on or before this date, as YYYY-MM-DD. Leave empty for no ceiling. |
workType | string | no | 'journal-article' | Restrict to one Crossref work type. Leave empty to match every type. |
publisherName | string | no | '—' | Substring match against the publisher name, e.g. Elsevier. Leave empty for any publisher. |
hasOrcid | boolean | no | False | Only return works where at least one author has an ORCID on file. |
hasAbstract | boolean | no | False | Only return works that carry an abstract. |
hasFullText | boolean | no | False | Only return works Crossref has a full-text link for. |
funderName | string | no | '—' | A funder's common name (e.g. Wellcome Trust) matched case-insensitively against each work's funder list,… |
sortField | string | no | 'score' | Field Crossref sorts results by. |
sortOrder | string | no | 'desc' | Ascending or descending. |
maxResults | integer | no | 100 | Stop after this many works. Each work is one billed result row. |
proxyConfiguration | object | no | {'useApifyProxy': False} | Crossref is a public API and does not need a proxy. Leave this off unless your account requires egress through Apify… |
Example input
{"query": "climate change adaptation","workType": "journal-article","maxResults": 3,"proxyConfiguration": {"useApifyProxy": false}}
📤 Output
Every row is one dataset item.
| Field | Type | Notes |
|---|---|---|
doi | string | Bare DOI string, e.g. 10.1037/xyz (no https://doi.org/ prefix). |
title | string | Work title. |
type | string | Crossref work type, e.g. journal-article. |
publisher | string | Publisher name. |
container_title | string | Journal or series title (first entry if Crossref lists several). |
authors | array | Author list, each with name, orcid, and affiliation. |
author_count | integer | Number of authors — convenience scalar. |
published_date | string | Best-available publication date (YYYY, YYYY-MM, or YYYY-MM-DD). |
volume | string | Journal volume. |
issue | string | Journal issue. |
page | string | Page range. |
issn | array | ISSNs (print and/or electronic). |
license_url | string | License URL (first entry if Crossref lists several). |
license_count | integer | Number of license entries Crossref returned. |
funders | array | Funder list, each with name, doi, and award_numbers. |
is_referenced_by_count | integer | Citation count. |
abstract | string | Abstract, passed through verbatim (JATS XML tags included, not stripped). |
url | string | Work URL — Crossref's own URL, or https://doi.org/<doi> when absent. |
Example output
{"doi": "10.1371/journal.pone.0212345","title": "Example study on citation graphs","type": "journal-article","publisher": "Public Library of Science (PLoS)","container_title": "PLOS ONE","authors": [{"name": "Jane Q. Researcher","orcid": "0000-0002-1234-5678","affiliation": "University of Somewhere"}],"author_count": 1,"published_date": "2023-04-12","volume": "18","issue": "4","page": "e0212345","issn": ["1932-6203"],"license_url": "https://creativecommons.org/licenses/by/4.0/","license_count": 1,"funders": [{"name": "Wellcome Trust","doi": "10.13039/100004440","award_numbers": ["WT12345"]}],"is_referenced_by_count": 42,"abstract": "<jats:p>We present an analysis of...</jats:p>","url": "https://doi.org/10.1371/journal.pone.0212345"}
💰 Pricing
Pay-Per-Event — you pay only when these events fire:
| Event | USD | What it is |
|---|---|---|
actor-start | $0.05 | One-off warm-up charge per run |
result | $0.002 | Per dataset item |
Example: 1 000 results at the rates above ≈ $2.05. No subscription, no minimum, no card to start — Apify gives every new account $5 of free credit.
🚧 Limitations
- Metadata only — it does not download PDFs or full text.
- The
abstractfield is passed through verbatim, JATS XML tags included; the Actor does not strip or reformat it. - Free-text funder-name matching is a client-side substring match and can return false positives (e.g. "Gates" matches multiple foundations) — pass the exact Funder DOI for precision.
❓ FAQ
Do I need an API key?
No. Crossref's works API is free and keyless. The Actor sends a polite, identifying User-Agent so runs stay inside Crossref's shared rate limit.
How deep can I page?
Crossref's cursor paging has no ~10 000-record wall the way offset paging does — the only limit is the Maximum works you set.
Can I filter by funder name instead of a Funder DOI?
Yes — type a funder's common name (e.g. "Wellcome Trust") and the Actor keeps only works whose funder list contains a case-insensitive match. Pass the exact Funder Registry DOI (e.g. 10.13039/100004440) for an exact server-side filter instead.
💬 Your feedback
Spotted a bug, hit a weird edge case, or need a new field? Open an issue on the Actor's Issues tab on Apify Console — we ship fixes weekly and we read every report.