dataset-recency-evidence-gate avatar

dataset-recency-evidence-gate

Pricing

from $10.00 / 1,000 results

Go to Apify Store
dataset-recency-evidence-gate

dataset-recency-evidence-gate

Check saved dataset timestamps before RAG ingestion. Evaluate up to 1,000 metadata records; get batch pass/fail and per-record diagnostics. No crawling or truth verification. n8n tutorial: https://recency-n8n-guide.saimislam.chatgpt.site

Pricing

from $10.00 / 1,000 results

Rating

0.0

(0)

Developer

Saim Islam

Saim Islam

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

a day ago

Last modified

Categories

Share

Dataset Recency & Evidence Gate

Check saved dataset metadata before ingestion. Supply an evaluation time, an age policy, and up to 1,000 records. Receive a batch decision and a diagnostic for each record. No website is crawled and no model API is called.

The Actor checks observation age separately from optional source-content age. Downloading an old document today does not make its contents newly updated. Only enable the source-age policy where content age matters to your workflow.

A pass means that the supplied metadata meets your policy. It does not verify timestamps, evidence authenticity, provenance, or factual truth.

Example

{
"as_of": "2026-09-24T12:00:00Z",
"max_age_seconds": 86400,
"max_source_age_seconds": 2592000,
"records": [
{
"id": "healthy",
"observed_at": "2026-09-24T11:00:00Z",
"source_updated_at": "2026-09-23T12:00:00Z",
"source_id": "feed-a",
"evidence_id": "capture-1"
},
{
"id": "old-observation",
"observed_at": "2026-09-22T12:00:00Z",
"source_updated_at": "2026-09-21T12:00:00Z",
"source_id": "feed-a",
"evidence_id": "capture-2"
},
{
"id": "fresh-fetch-old-content",
"observed_at": "2026-09-24T11:00:00Z",
"source_updated_at": "2026-08-01T12:00:00Z",
"source_id": "feed-a",
"evidence_id": "capture-3"
},
{
"id": "unknown-source-age",
"observed_at": "2026-09-24T11:00:00Z",
"source_updated_at": null,
"source_id": "feed-a",
"evidence_id": "capture-4"
}
]
}

This fixed historical example always returns decision: "fail": one record passes and three are blocked. Its codes are STALE, SOURCE_STALE, and SOURCE_TIMESTAMP_MISSING. To get a passing example, keep only the healthy record. For real work, provide your intended evaluation time and saved metadata; the Actor never silently substitutes the current clock.

Input

FieldMeaning
as_ofRequired evaluation timestamp with an explicit timezone.
max_age_secondsRequired inclusive observation-age limit, an integer from 0 to 315,360,000.
max_source_age_secondsOptional content-age limit with the same range. When present, every record needs a source timestamp. Omit the field to disable this check.
records1–1,000 objects with unique, nonblank id values of at most 128 characters and no surrounding whitespace.
record.observed_atDeclared time this exact source observation was captured.
record.source_updated_atDeclared content-update time represented in that observation. Requires an explicit source-age policy.
record.source_idNonblank source identifier, at most 512 characters. No source is fetched.
record.evidence_idNonblank reference to the observation, at most 512 characters. No evidence is authenticated.

Accepted timestamp forms are YYYY-MM-DDTHH:MM:SSZ and a numeric timezone offset, optionally with one to six fractional second digits. Date-only, timezone-free, unknown -00:00 offsets, invalid calendar dates, and leap seconds are rejected. Offsets are normalized to UTC. Equality with an age limit passes; future times, even by one microsecond, fail. Clock-skew tolerance is not implemented.

Missing or invalid row timestamps and evidence fields block that row. Source updates later than observation time are inconsistent and also block it. Malformed policies, unknown fields, duplicate IDs, or invalid batch sizes fail the run before scoring. Supplying source_updated_at without a source-age policy is an input error, even when the value is null. No rows are silently dropped.

Normalized input JSON is limited to 1,000,000 UTF-8 bytes. Raw documents and page bodies are outside this contract. The platform parses JSON before the evaluator; the Actor does not guarantee detection of duplicate JSON object keys.

Output

The default dataset contains one item for the whole batch, with:

  • decision: pass only when every record meets the policy.
  • summary: total, passed, and blocked record counts plus diagnostic counts.
  • results: record IDs, ages in seconds, recency statuses, evidence completeness, decisions, and diagnostic codes.
  • policy and normalized as_of: the exact policy and evaluation clock used.
  • provenance_verified: false and truth_verified: false.

The Summary view shows the batch counts. The Record diagnostics view contains the nested results array and policy. Download JSON for the complete report.

A completed evaluation succeeds as an Actor run even when its policy decision is fail. API and automation clients must inspect decision in the output; successful run status alone does not mean records passed. Invalid batch input or a storage failure fails the Actor run.

Recency status is fresh, stale, unknown, or invalid. A disabled source-age check is not_checked. Missing or invalid times have null ages; future times may have negative ages. Multiple diagnostics can apply to one record, so diagnostic counts can exceed the number of blocked records.

Using saved crawler metadata

For a saved Website Content Crawler dataset, prepare input with the following mapping:

Saved fieldGate field
crawl.loadedTimeobserved_at
crawl.loadedUrl, falling back to urlsource_id
Your dataset ID plus the item's original offsetevidence_id, such as dataset:ID:item:20
An opaque unique row IDid

Preserve the original dataset order and offset when constructing evidence references. These references remain declarations; the Actor does not open the dataset to verify them. Supply as_of and max_age_seconds yourself.

Do not map crawl.loadedTime to source_updated_at. A load timestamp says when the page was fetched. If reliable source-update metadata is unavailable, omit the source-age policy or accept an explicit missing-source-time failure when that policy is required. This Actor accepts prepared metadata; automatic dataset downloading is not implemented.

Data and cost

Apify receives and stores run input and output under its storage settings. Input includes the source and evidence identifiers you supply. Output contains opaque record IDs, policy metadata, and diagnostics; it does not echo source or evidence identifiers. Use opaque IDs and avoid secrets in all input fields. No automatic deletion or zero-retention behavior is promised.

There are no third-party model charges. Consult the current pricing shown before running for platform execution/storage charges and any Actor fee. Local timing is not evidence of cloud cost.

Memory is configured at 256 MB. The process has a 60-second deadline beginning when the Python entrypoint starts, including SDK imports and storage operations. Container startup is additional. Partial output is not guaranteed on timeout. These limits bound individual work, not account spending across unlimited runs.

Version 0.1.1 adds Apify packaging and the saved-crawler metadata example to the v0.1.0 deterministic evaluator. Provenance and truth verification remain absent.