Crossref Academic Paper Metadata Scraper avatar

Crossref Academic Paper Metadata Scraper

Pricing

from $3.62 / 1,000 results

Go to Apify Store
Crossref Academic Paper Metadata Scraper

Crossref Academic Paper Metadata Scraper

Scrapes academic paper metadata from Crossref by search query or publication type. Returns DOI, title, authors, journal, citation counts, and dates.

Pricing

from $3.62 / 1,000 results

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

ParseForge

Crossref Academic Paper Metadata Scraper

Scrape Crossref academic paper metadata for any DOI, search query, or publication type, up to a million records per run. Each record includes title, authors, journal, DOI, citation counts, and publication dates. No API key required. Export to CSV, JSON, Excel, or XML.

The Actor queries the public Crossref REST API, filters by publication type, search query, and sort order, and returns each matching scholarly work as one flat row. It covers journal articles, books, dissertations, datasets, and more.

Who uses itWhat they scrape Crossref for
Academic researchersBuilding a bibliography for a literature review
LibrariansVerifying citation metadata for institutional repositories
Data scientistsGathering a corpus of paper metadata for analysis
Journal editorsChecking DOI registration and metadata completeness
PhD studentsCollecting references for a thesis

What it does

This Actor collects academic paper metadata from Crossref by search query or publication type, and returns each work as a flat row with DOI, title, authors, journal, citation counts, and dates.

  • 🔎 Search query: free-text search across titles and other metadata fields.
  • 📚 Publication type filter: journal-article, book-chapter, book, proceedings-article, dissertation, report, standard, dataset, or posted-content.
  • 📅 Sort order: by publication date or relevance to the search query.
  • 📦 Bulk export: up to 1,000,000 records per run for paid users.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

What you can do with Crossref data

📚 Build a literature review bibliography.

A researcher enters a search query like "climate change migration" and exports all matching journal articles with full metadata to CSV for reference management.

🔍 Verify DOI metadata for a journal issue.

An editor filters by journal-article and searches for a specific title to confirm authors, volume, issue, and page numbers before publication.

📊 Analyze publication trends.

A data scientist fetches all journal articles sorted by published date for a given year and uses citation counts to identify influential papers.

🎓 Collect references for a thesis.

A PhD student searches for a topic, filters to dissertations and journal articles, and exports the metadata to build a reference list.

Why choose this scraper

What you get
No API keyUses the public Crossref REST API without authentication
Rich metadataReturns DOI, title, authors, journal, citation counts, and dates
Flexible filteringSearch by text or filter by scholarly work type
ScalableFetch up to a million records per run

How it compares

This Actor focuses on flexible search and filtering of Crossref metadata with no API key required, similar to other Crossref scrapers but with a simpler input schema.

FeatureParseForgeCrossRef Academic Metadata ScraperCrossref Academic Paper SearchCrossref Api Scraper
Search by text queryYesYesYesYes
Filter by publication typeYesNot listedNot listedYes
Sort by relevanceYesNot listedNot listedNot listed
Fetch up to 1,000,000 recordsYesNot listedNot listedNot listed
No API key requiredYesNot listedNot listedYes

What a Crossref record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

{
"DOI": "10.1157/13053466",
"type": "journal-article",
"title": "Comentario: Prevención de los factores de riesgo de los trastornos de la conducta alimentaria en adolescentes:",
"containerTitle": "Atención Primaria",
"shortContainerTitle": "Aten Primaria",
"publisher": "Elsevier BV",
"issue": "7",
"volume": "32",
"page": "408-409",
"publishedPrintDate": "2203-10",
"issuedDate": "2203-10",
"indexedDateTime": "2025-05-30T05:22:24Z",
"indexedTimestamp": 1748582544347,
"createdDateTime": "2003-11-17T16:44:58Z",
"createdTimestamp": 1069087498000
}

Every value above comes from a real run. A field a record does not have comes back as null.

Configure the run

Drive the Actor with a search query or leave it empty to fetch the most recently published journal articles. Filter by publication type and sort by published date or relevance. The Input tab lists every parameter.

A first run with the defaults:

{
"maxItems": 10,
"filterType": "journal-article",
"sortOrder": "published"
}

A larger pull:

{
"maxItems": 200,
"filterType": "journal-article",
"sortOrder": "published"
}

Free users

Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.

Run it

  1. Create a free Apify account.
  2. Open the Crossref Academic Paper Metadata Scraper.
  3. Set your inputs and any filters, then click Start.
  4. Export the results as CSV, Excel, JSON, or XML from the Dataset tab.

Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.

Use with AI agents (MCP)

Give an AI agent live access to Crossref through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

$claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/crossref-academic-paper-scraper"

Then prompt it in plain language to run the scraper and read back the results.

Troubleshooting

Why am I getting no results?

Check your search query for typos or overly specific terms. Try a broader query or leave the search field empty to fetch recent articles.

Why are some fields empty?

Crossref metadata depends on what publishers deposit. Some fields may be missing for certain records.

Why is my run limited to 10 items?

Free users are limited to 10 items as a preview. Upgrade to a paid plan to fetch up to 1,000,000 records.

Why does sorting by relevance not work?

Relevance sorting requires a search query. If the query is empty, results are sorted by publication date.

FAQ

QuestionAnswer
Do I need a Crossref API key?No, this Actor uses the public Crossref REST API without authentication. You can start scraping immediately.
What metadata fields are returned?Each record includes DOI, type, title, containerTitle, shortContainerTitle, publisher, issue, volume, page, publishedPrintDate, issuedDate, indexedDateTime, indexedTimestamp, createdDateTime, createdTimestamp, depositedDateTime, depositedTimestamp, isReferencedByCount, referencesCount, URL, ISSN, issnPrint, issnElectronic, licenseUrl, licenseContentVersion, licenseDelayInDays, linkUrl, linkContentType, linkContentVersion, linkIntendedApplication, resourcePrimaryUrl, source, member, prefix, score, alternativeId, journalIssue, authors, scrapedAt, error.
Can I search by DOI?Yes, you can enter a DOI as a search query to retrieve metadata for a specific paper.
What publication types are supported?Journal articles, book chapters, books, proceedings articles, dissertations, reports, standards, datasets, and posted content.
How many records can I fetch?Free users are limited to 10 items as a preview. Paid users can fetch up to 1,000,000 records per run.
Can I sort results?Yes, sort by publication date (newest first) or by relevance to your search query.
Is the data from Crossref complete?Crossref contains metadata for over 150 million scholarly works, but some records may have missing fields depending on what publishers deposited.
Can I export to Excel?Yes, you can export results to CSV, JSON, Excel, or XML.
Does this Actor get abstracts?The dataset does not include an abstracts field.
How do I cite the data?You should cite the original papers, not the metadata. Crossref metadata is provided under a CC0 license.

Browse the full ParseForge collection for more scrapers.

🆘 Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Crossref. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

💰 How much does it cost to scrape Crossref Academic Paper Metadata?

This Actor uses pay-per-result pricing: $0.004 per result collected. You are billed only for the results you receive, so a run that returns nothing costs nothing.