CKAN Resource Link Change Monitor
Pricing
$0.01 / successful manifest check
CKAN Resource Link Change Monitor
Track CKAN resource download URL, format, MIME type, hash, size and DataStore metadata changes. Flag missing required resource IDs for open-data pipelines. Failed checks preserve history.
Pricing
$0.01 / successful manifest check
Rating
0.0
(0)
Developer
QuietDataTools
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Keep CKAN download manifests current. Monitor one public dataset for resource URL, format, MIME type, name, reported hash/size and DataStore availability changes. Stable resource IDs distinguish a renamed or moved download from replacement resources. Exact required resource IDs flag a missing input to your pipeline.
Useful for open-data ingestion, catalog maintenance and scheduled ETL preflight. It reads package metadata only and never downloads the resource files or dataset rows.
Quick start
{"portalUrl":"https://open.canada.ca/data/en","datasetId":"09ffaeb5-ec8f-5bb5-bdcb-3436ccf26f58","monitorName":"default","requiredResourceIds":[],"includeUnchanged":true}
The example is Canada's Climatic Regions dataset. Other HTTPS CKAN portals and portal subpaths are supported if they expose the standard public api/3/action/package_show endpoint. This is CKAN only; data.gov's successor API, Socrata and DCAT-only catalogs are different protocols.
First check returns BASELINE. Later checks return ADDED, CHANGED, UNCHANGED or NO_LONGER_LISTED. Each row includes exact previous/current metadata, changed fields, source URL and check time. NO_LONGER_LISTED means absent from a successfully read manifest; it does not prove the underlying file was deleted. The OUTPUT summary includes resource counts, missing required IDs and the prior check time. Suppress unchanged rows without suppressing successful check summaries.
Pricing
$0.01 per successful manifest check, including baseline, empty manifest and unchanged checks. Platform usage is included. Failed source/validation checks do not emit the successful-check charge and preserve the last good snapshot. One dataset per run. Use an Apify schedule and consume the default dataset/OUTPUT in your workflow; the Actor does not send messages itself.
Boundaries
- Public HTTPS metadata only; no credentials, cookies, custom headers, proxies or gated datasets. DNS is checked for public addresses and pinned for each request; only same-origin redirects are permitted.
- 2 MB, 20-second source deadline, at most 1,000 resources and 100 required IDs. Incomplete, duplicate, mismatched or inactive source records fail rather than inventing removals.
- Resource URLs and hashes are source-reported strings. No link reachability, byte integrity, license, freshness or downstream compatibility certification. Empty/missing hash and format are kept distinct. Timestamp-only updates are ignored.
- History is isolated by user, Actor, portal, requested dataset and monitor name. A slug retargeting to another dataset ID fails; choose a new monitor name deliberately. A new input dataset/URL is a separate history.
- Schedule one concurrent run per monitor. The lease is best effort and non-atomic. Interrupted delivery/billing/history writes can repeat output or charges on retry (at-least-once behavior).
- Portal extensions can alter field shapes. Unsupported metadata fails safely. Check the dataset's own license and source terms before using its data.
The Apify SDK retains a transitive http-cache-semantics advisory with no upstream patch at release. The Actor's source client uses native HTTPS, no credentials and no application shared cache; this does not claim the SDK dependency is cleared. Source and fixture tests are retained privately in the cloud project.