DataCite Scraper (Datasets, Software, Research Outputs) avatar

DataCite Scraper (Datasets, Software, Research Outputs)

Pricing

from $1.40 / 1,000 results

Go to Apify Store
DataCite Scraper (Datasets, Software, Research Outputs)

DataCite Scraper (Datasets, Software, Research Outputs)

Search DataCite, the CC0 DOI registry for research outputs beyond journal articles -- 133M+ datasets, software, samples, instruments, images, text, workflows and dissertations. Lucene query plus filters (type, publisher, year, provider, ORCID, ROR, funding). Cursor pagination past 10k. DOI lookup.

Pricing

from $1.40 / 1,000 results

Rating

0.0

(0)

Developer

Ibnu Adzim

Ibnu Adzim

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 days ago

Last modified

Share

Search DataCite — the DOI registry for the research outputs that aren't journal articles: datasets, software, samples, instruments, images, text, workflows, computational notebooks, dissertations, preprints. ~133M DOIs, metadata CC0, no key, no login.

Companion to openalex-scholar-scraper (works), crossref-works-scraper (articles) and pubmed-articles-scraper (biomedical).

  • search — a Lucene query and/or structured filters (resource type, publisher / repository, publication year range, provider, client, DOI prefix, ORCID, ROR, subject, has-citations / views / downloads / funding). Cursor-paginated — no 10,000-result ceiling.
  • dois — look up specific DOIs → the full record each.
Record typeOne perCarries
SEARCH_SUMMARYsearch runcomposed query, upstreamCount, resultsReturned, cursorPages, filtersApplied
RESULTDOIdoi, title, resourceTypeGeneral, publisher, publicationYear, creators (+ ORCID + affiliation), description, subjects, citationCount / viewCount / downloadCount, license, geoLocations, fundingReferences, relatedIdentifiers, containerTitle
ERRORbad input / missing DOI_error + _errorDetail

Every RESULT row carries the verbatim DataCite record in raw (drop with slimOutput).

Things this API will mislead you about

Each is measured, and each has a scenario in tests/smoke/datacite-scraper_traps.sh (10/10 passing).

page[number] pagination is capped at 10,000 results. This Actor uses page[cursor] + links.next throughout, which has no ceiling — deepPaginationViaCursor on the summary confirms it.

An unknown resourceType returns total: 0, silently — a "no such output" that is really "no such type". resourceType is validated against DataCite's controlled list up front.

An empty query matches all 133M records. A query is not required, but a query or a filter is — a bare search is rejected.

page[size] maxes at 1000 (a larger value is silently clamped).

titles, creators, descriptions, subjects, rightsList are arrays of objects, often empty, and creators[].name can be "Anonymous". The normaliser takes the first title, joins creator names into creatorNames, and pulls the first Abstract-type description.

A DOI lookup is case-insensitive and the slash may be raw or %2F. A DOI that DataCite never registered (e.g. a Crossref journal DOI like 10.1038/nature12373) is an honest HTTP 404 → record_not_found with a hint to try crossref-works-scraper.

Notes on cost

search: one request per 1000 results (cursor pages). dois: one request per DOI. No proxy needed.