DataCite Scraper (Datasets, Software, Research Outputs)
Pricing
from $1.40 / 1,000 results
DataCite Scraper (Datasets, Software, Research Outputs)
Search DataCite, the CC0 DOI registry for research outputs beyond journal articles -- 133M+ datasets, software, samples, instruments, images, text, workflows and dissertations. Lucene query plus filters (type, publisher, year, provider, ORCID, ROR, funding). Cursor pagination past 10k. DOI lookup.
Pricing
from $1.40 / 1,000 results
Rating
0.0
(0)
Developer
Ibnu Adzim
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
6 days ago
Last modified
Categories
Share
Search DataCite — the DOI registry for the research outputs that aren't journal articles: datasets, software, samples, instruments, images, text, workflows, computational notebooks, dissertations, preprints. ~133M DOIs, metadata CC0, no key, no login.
Companion to openalex-scholar-scraper (works), crossref-works-scraper
(articles) and pubmed-articles-scraper (biomedical).
search— a Lucenequeryand/or structured filters (resource type, publisher / repository, publication year range, provider, client, DOI prefix, ORCID, ROR, subject, has-citations / views / downloads / funding). Cursor-paginated — no 10,000-result ceiling.dois— look up specific DOIs → the full record each.
| Record type | One per | Carries |
|---|---|---|
SEARCH_SUMMARY | search run | composed query, upstreamCount, resultsReturned, cursorPages, filtersApplied |
RESULT | DOI | doi, title, resourceTypeGeneral, publisher, publicationYear, creators (+ ORCID + affiliation), description, subjects, citationCount / viewCount / downloadCount, license, geoLocations, fundingReferences, relatedIdentifiers, containerTitle |
ERROR | bad input / missing DOI | _error + _errorDetail |
Every RESULT row carries the verbatim DataCite record in raw (drop with
slimOutput).
Things this API will mislead you about
Each is measured, and each has a scenario in
tests/smoke/datacite-scraper_traps.sh (10/10 passing).
page[number] pagination is capped at 10,000 results. This Actor uses
page[cursor] + links.next throughout, which has no ceiling —
deepPaginationViaCursor on the summary confirms it.
An unknown resourceType returns total: 0, silently — a "no such
output" that is really "no such type". resourceType is validated against
DataCite's controlled list up front.
An empty query matches all 133M records. A query is not required, but a
query or a filter is — a bare search is rejected.
page[size] maxes at 1000 (a larger value is silently clamped).
titles, creators, descriptions, subjects, rightsList are arrays
of objects, often empty, and creators[].name can be "Anonymous". The
normaliser takes the first title, joins creator names into creatorNames,
and pulls the first Abstract-type description.
A DOI lookup is case-insensitive and the slash may be raw or %2F. A DOI
that DataCite never registered (e.g. a Crossref journal DOI like
10.1038/nature12373) is an honest HTTP 404 → record_not_found with a hint
to try crossref-works-scraper.
Notes on cost
search: one request per 1000 results (cursor pages). dois: one request per
DOI. No proxy needed.