# Changelog of Crossref Scholarly Scraper — DOIs, Citations & Journals (`logiover/crossref-scraper`) Actor

- **URL**: https://apify.com/logiover/crossref-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/logiover/crossref-scraper.md

## Changelog

### 2026-09-01

- Fleet-wide health check. Verified against this Actor's real run history: 30-day success rate, output row counts, per-field fill rates, and peak memory against the configured memory limit.
- Reviewed for the failure patterns that have cost this fleet runs — unguarded proxy setup, retry loops that can outlast the run's own time budget, and full-page HTML parsing that can exhaust a small container.
- No change to input, output fields or scraping logic.

### 2026-08-11

- Maintenance release: refreshed the build and dependencies.
- Re-verified live execution, non-empty structured output and dataset field/type integrity.
- Reviewed reliability (retries, pagination) and output quality as part of a full-fleet QA pass.

### 2026-08-01

- Completed the August 2026 full health check: verified empty/programmatic default, Console UI default, and two source-informed alternative inputs on Apify.
- Confirmed successful live execution, non-empty structured output, dataset-field/type integrity, and logical sample quality within the 5-minute quality window.

### 2026-07-18

- Maintenance & reliability pass — re-verified end-to-end against live data via the Apify API; confirmed the Actor completes successfully and returns non-empty, well-formed results within the 5-minute quality window on its default input.
- Refreshed build so the latest version is current; no change to inputs, output fields or scraping logic.

### 2026-07-16

- Right-sized memory allocation to cut per-run compute cost: capped RAM at 1 GB (was defaulting to 4 GB) for this lightweight JSON-API scraper. No logic or output changes.

### 2026-07-12

- **Empty input now returns data (was failing before).** Previously an empty `{}` run hard-failed with exit code 1 because `works` mode required a query/filter. Now empty input browses the most-cited works (`sort=is-referenced-by-count`), bounded by `maxResults`. All modes fall back to this browse behavior when their required input (issn / dois) is missing, instead of failing.
- **No mandatory fields.** Removed `required: ["mode"]` → `required: []`.
- **New `type` dropdown** (select) with all 30 Crossref work types (journal article, book chapter, dataset, dissertation, etc.). Merged into the Crossref filter automatically.
- **New `sort` and `order` dropdowns** (most cited / publication date / recently indexed / relevance; descending or ascending).
- **Numbers are real numbers.** `citationCount` and `referencesCount` are now numeric instead of strings. Added `publishedYear` as a numeric field.
- **AUTO proxy — now actually used.** The proxy input was previously ignored; requests now go through automatic Apify Proxy with a direct-connection fallback on the final retry (Crossref is a public API).
- **Bounded & graceful.** Long paginations stop after ~4 minutes and exit successfully with whatever was collected.
- Added per-field `fields` JSON-schema (titles + descriptions) to the dataset schema.

Note: existing output field keys are unchanged; `publishedYear` was added.
