dataset-recency-evidence-gate
Pricing
from $10.00 / 1,000 results
dataset-recency-evidence-gate
Check saved dataset timestamps before RAG ingestion. Evaluate up to 1,000 metadata records; get batch pass/fail and per-record diagnostics. No crawling or truth verification. n8n tutorial: https://recency-n8n-guide.saimislam.chatgpt.site
Pricing
from $10.00 / 1,000 results
Rating
0.0
(0)
Developer
Saim Islam
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
a day ago
Last modified
Categories
Share
Dataset Recency & Evidence Gate
Check saved dataset metadata before ingestion. Supply an evaluation time, an age policy, and up to 1,000 records. Receive a batch decision and a diagnostic for each record. No website is crawled and no model API is called.
The Actor checks observation age separately from optional source-content age. Downloading an old document today does not make its contents newly updated. Only enable the source-age policy where content age matters to your workflow.
A pass means that the supplied metadata meets your policy. It does not verify timestamps, evidence authenticity, provenance, or factual truth.
Example
{"as_of": "2026-09-24T12:00:00Z","max_age_seconds": 86400,"max_source_age_seconds": 2592000,"records": [{"id": "healthy","observed_at": "2026-09-24T11:00:00Z","source_updated_at": "2026-09-23T12:00:00Z","source_id": "feed-a","evidence_id": "capture-1"},{"id": "old-observation","observed_at": "2026-09-22T12:00:00Z","source_updated_at": "2026-09-21T12:00:00Z","source_id": "feed-a","evidence_id": "capture-2"},{"id": "fresh-fetch-old-content","observed_at": "2026-09-24T11:00:00Z","source_updated_at": "2026-08-01T12:00:00Z","source_id": "feed-a","evidence_id": "capture-3"},{"id": "unknown-source-age","observed_at": "2026-09-24T11:00:00Z","source_updated_at": null,"source_id": "feed-a","evidence_id": "capture-4"}]}
This fixed historical example always returns decision: "fail": one record passes
and three are blocked. Its codes are STALE, SOURCE_STALE, and
SOURCE_TIMESTAMP_MISSING. To get a passing example, keep only the healthy
record. For real work, provide your intended evaluation time and saved metadata;
the Actor never silently substitutes the current clock.
Input
| Field | Meaning |
|---|---|
as_of | Required evaluation timestamp with an explicit timezone. |
max_age_seconds | Required inclusive observation-age limit, an integer from 0 to 315,360,000. |
max_source_age_seconds | Optional content-age limit with the same range. When present, every record needs a source timestamp. Omit the field to disable this check. |
records | 1–1,000 objects with unique, nonblank id values of at most 128 characters and no surrounding whitespace. |
record.observed_at | Declared time this exact source observation was captured. |
record.source_updated_at | Declared content-update time represented in that observation. Requires an explicit source-age policy. |
record.source_id | Nonblank source identifier, at most 512 characters. No source is fetched. |
record.evidence_id | Nonblank reference to the observation, at most 512 characters. No evidence is authenticated. |
Accepted timestamp forms are YYYY-MM-DDTHH:MM:SSZ and a numeric timezone offset,
optionally with one to six fractional second digits. Date-only, timezone-free,
unknown -00:00 offsets, invalid calendar dates, and leap seconds are rejected.
Offsets are normalized to UTC. Equality with an age limit passes; future times,
even by one microsecond, fail. Clock-skew tolerance is not implemented.
Missing or invalid row timestamps and evidence fields block that row. Source
updates later than observation time are inconsistent and also block it. Malformed
policies, unknown fields, duplicate IDs, or invalid batch sizes fail the run before
scoring. Supplying source_updated_at without a source-age policy is an input
error, even when the value is null. No rows are silently dropped.
Normalized input JSON is limited to 1,000,000 UTF-8 bytes. Raw documents and page bodies are outside this contract. The platform parses JSON before the evaluator; the Actor does not guarantee detection of duplicate JSON object keys.
Output
The default dataset contains one item for the whole batch, with:
decision:passonly when every record meets the policy.summary: total, passed, and blocked record counts plus diagnostic counts.results: record IDs, ages in seconds, recency statuses, evidence completeness, decisions, and diagnostic codes.policyand normalizedas_of: the exact policy and evaluation clock used.provenance_verified: falseandtruth_verified: false.
The Summary view shows the batch counts. The Record diagnostics view contains the
nested results array and policy. Download JSON for the complete report.
A completed evaluation succeeds as an Actor run even when its policy decision
is fail. API and automation clients must inspect decision in the output;
successful run status alone does not mean records passed. Invalid batch input or
a storage failure fails the Actor run.
Recency status is fresh, stale, unknown, or invalid. A disabled source-age
check is not_checked. Missing or invalid times have null ages; future times may
have negative ages. Multiple diagnostics can apply to one record, so diagnostic
counts can exceed the number of blocked records.
Using saved crawler metadata
For a saved Website Content Crawler dataset, prepare input with the following mapping:
| Saved field | Gate field |
|---|---|
crawl.loadedTime | observed_at |
crawl.loadedUrl, falling back to url | source_id |
| Your dataset ID plus the item's original offset | evidence_id, such as dataset:ID:item:20 |
| An opaque unique row ID | id |
Preserve the original dataset order and offset when constructing evidence
references. These references remain declarations; the Actor does not open the
dataset to verify them. Supply as_of and max_age_seconds yourself.
Do not map crawl.loadedTime to source_updated_at. A load timestamp says when
the page was fetched. If reliable source-update metadata is unavailable, omit the
source-age policy or accept an explicit missing-source-time failure when that
policy is required. This Actor accepts prepared metadata; automatic dataset
downloading is not implemented.
Data and cost
Apify receives and stores run input and output under its storage settings. Input includes the source and evidence identifiers you supply. Output contains opaque record IDs, policy metadata, and diagnostics; it does not echo source or evidence identifiers. Use opaque IDs and avoid secrets in all input fields. No automatic deletion or zero-retention behavior is promised.
There are no third-party model charges. Consult the current pricing shown before running for platform execution/storage charges and any Actor fee. Local timing is not evidence of cloud cost.
Memory is configured at 256 MB. The process has a 60-second deadline beginning when the Python entrypoint starts, including SDK imports and storage operations. Container startup is additional. Partial output is not guaranteed on timeout. These limits bound individual work, not account spending across unlimited runs.
Version 0.1.1 adds Apify packaging and the saved-crawler metadata example to the v0.1.0 deterministic evaluator. Provenance and truth verification remain absent.