DataCite Metadata Scraper avatar

DataCite Metadata Scraper

Pricing

Pay per event

Go to Apify Store
DataCite Metadata Scraper

DataCite Metadata Scraper

Comprehensive DataCite metadata scraper for extracting DOI metadata from DataCite API. Perfect for researchers, librarians, and data scientists who need automated access to scholarly publication metadata, research datasets, and digital object identifiers.

Pricing

Pay per event

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

16 hours ago

Last modified

Share

ParseForge Banner

๐Ÿ“š DataCite Metadata Scraper

๐Ÿš€ Collect scholarly DOI metadata and research dataset records in seconds. Filter by keyword, repository, publisher, resource type, and year. No coding, no DataCite account required.

Pull structured records from DataCite Metadata โ€” clean fields ready as CSV, JSON, JSONL, Excel, or XML for downstream pipelines.

Copy to your AI assistant

Copy this block into ChatGPT, Claude, Cursor, or any LLM to start using this actor.

parseforge/datacite-metadata-scraper on Apify. Call: ApifyClient("TOKEN").actor("parseforge/datacite-metadata-scraper").call(run_input={...}), then client.dataset(run["defaultDatasetId"]).list_items().items for results. Key inputs: maxItems (integer, default 10), query (string, default "climate"), doi (string), repositoryId (string), publisher (string), resourceType (string). Full actor spec: fetch build via GET https://api.apify.com/v2/acts/parseforge~datacite-metadata-scraper (Bearer TOKEN). Get token: https://console.apify.com/account/integrations

The DataCite Metadata Scraper retrieves Digital Object Identifier (DOI) metadata from the DataCite registry, which indexes over 45 million DOIs across academic publications, research datasets, software, and other scholarly outputs. Each record includes the DOI, title, publisher, publication year, resource type, creation date, update date, and a resolvable URL. You can filter by keyword, specific DOI, repository (Zenodo, Dryad, Figshare, Dataverse), publisher, resource type, and publication year. Free users can collect up to 10 records per run, while paid users can retrieve up to 1,000,000.

Whether you are building a literature database for a systematic review, analyzing publication trends across institutions, tracking open data availability in your research field, or monitoring repository output over time, this tool replaces hours of manual DOI lookups with a single automated query. Results export to JSON, CSV, or Excel for immediate use in citation managers, bibliometric tools, or data analysis pipelines. The scraper handles pagination and rate limiting automatically, letting you focus on research instead of data collection.

Target AudienceUse Cases
Academic ResearchersBuild literature databases and track publications in specific fields
Research LibrariansCatalog DOI records and monitor repository output
Data ScientistsAnalyze publication trends and research metadata at scale
Institutional AnalystsTrack publication volume and output across departments
Science Policy AnalystsStudy open data availability and repository growth
Bibliometric ResearchersCollect DOI metadata for citation and impact analysis

๐Ÿ“‹ What the DataCite Metadata Scraper does

  • ๐Ÿ“š DOI records - retrieve the full Digital Object Identifier for each scholarly output, ready for citation or resolution
  • ๐Ÿท๏ธ Titles - extract publication or dataset titles for cataloging and search
  • ๐Ÿ“ฐ Publishers - capture the organization or institution that registered the DOI
  • ๐Ÿ“… Publication years - filter and sort by year to focus on recent research or historical trends
  • ๐Ÿ—‚๏ธ Resource types - classify records as datasets, articles, software, images, or other scholarly object types
  • ๐Ÿ”— Resolvable URLs - get working DOI links that resolve to the full publication or dataset landing page

The scraper queries the DataCite REST API and iterates through paginated results using your specified filters. Each record is normalized with consistent field names and pushed to an Apify dataset in real time. You can look up a single DOI or search across the entire DataCite registry with keyword and faceted filters.

๐Ÿ’ก Why it matters: DataCite indexes DOIs from over 2,000 data centers worldwide. Manually searching and downloading metadata is tedious. This scraper gives you structured, filterable access to the registry in minutes.

๐Ÿ“Š Data fields

Each record includes: container, contributors, created, creators, dates, descriptions, doi, doiUrl, formats, fundingReferences, geoLocations, language, publicationYear, publisher, registered, relatedIdentifiers, resourceType, resourceTypeGeneral, rightsList, schemaVersion, scrapedTimestamp, sizes, subjects, title, updated, url, version. All 27 field names come from a real production run, so what you see here is what lands in your dataset.

โš ๏ธ Good to Know: Free users are automatically limited to 10 items per run. When a specific DOI is provided, only that single record is returned. Leave the query field empty to browse all records with other filters applied.

๐Ÿš€ How to use

  1. Sign up - Create a free Apify account with $5 credit
  2. Find the Actor - Search for "DataCite Metadata Scraper" in the Apify Store
  3. Set your search criteria - Enter keywords, resource type, year, or a specific DOI
  4. Start the run - Click "Start" and watch results appear in real time
  5. Export your data - Download as JSON, CSV, or Excel from the dataset tab

๐Ÿ•’ Typical run time: 15 to 60 seconds for up to 100 records. Larger runs with 1,000+ records may take a few minutes depending on the query scope.

ActorDescription
Hugging Face Model ScraperCollect model metadata and download stats from Hugging Face
PR Newswire ScraperCollect press releases and research announcements
GSA eLibrary ScraperCollect government contractor and vendor data
Greatschools ScraperExtract school ratings and performance data
Smart Apify Actor ScraperScrape Apify actor metadata with 70+ fields

๐Ÿ’ก Pro Tip: Combine the DataCite Metadata Scraper with the Hugging Face Model Scraper to cross-reference published datasets with ML models trained on them.

Disclaimer: This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by DataCite, Zenodo, Dryad, Figshare, or any data center. All trademarks mentioned are the property of their respective owners.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.