# dataset-recency-evidence-gate (`saimislam/dataset-recency-evidence-gate`) Actor

Check saved dataset timestamps before RAG ingestion. Evaluate up to 1,000 metadata records; get batch pass/fail and per-record diagnostics. No crawling or truth verification. n8n tutorial: https://recency-n8n-guide.saimislam.chatgpt.site

- **URL**: https://apify.com/saimislam/dataset-recency-evidence-gate.md
- **Developed by:** [Saim Islam](https://apify.com/saimislam) (community)
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Dataset Recency & Evidence Gate

Check saved dataset metadata before ingestion. Supply an evaluation time, an age
policy, and up to 1,000 records. Receive a batch decision and a diagnostic for each
record. No website is crawled and no model API is called.

The Actor checks **observation age** separately from optional **source-content
age**. Downloading an old document today does not make its contents newly updated.
Only enable the source-age policy where content age matters to your workflow.

**A pass means that the supplied metadata meets your policy. It does not verify
timestamps, evidence authenticity, provenance, or factual truth.**

### Example

```json
{
  "as_of": "2026-09-24T12:00:00Z",
  "max_age_seconds": 86400,
  "max_source_age_seconds": 2592000,
  "records": [
    {
      "id": "healthy",
      "observed_at": "2026-09-24T11:00:00Z",
      "source_updated_at": "2026-09-23T12:00:00Z",
      "source_id": "feed-a",
      "evidence_id": "capture-1"
    },
    {
      "id": "old-observation",
      "observed_at": "2026-09-22T12:00:00Z",
      "source_updated_at": "2026-09-21T12:00:00Z",
      "source_id": "feed-a",
      "evidence_id": "capture-2"
    },
    {
      "id": "fresh-fetch-old-content",
      "observed_at": "2026-09-24T11:00:00Z",
      "source_updated_at": "2026-08-01T12:00:00Z",
      "source_id": "feed-a",
      "evidence_id": "capture-3"
    },
    {
      "id": "unknown-source-age",
      "observed_at": "2026-09-24T11:00:00Z",
      "source_updated_at": null,
      "source_id": "feed-a",
      "evidence_id": "capture-4"
    }
  ]
}
```

This fixed historical example always returns `decision: "fail"`: one record passes
and three are blocked. Its codes are `STALE`, `SOURCE_STALE`, and
`SOURCE_TIMESTAMP_MISSING`. To get a passing example, keep only the `healthy`
record. For real work, provide your intended evaluation time and saved metadata;
the Actor never silently substitutes the current clock.

### Input

| Field | Meaning |
| --- | --- |
| `as_of` | Required evaluation timestamp with an explicit timezone. |
| `max_age_seconds` | Required inclusive observation-age limit, an integer from 0 to 315,360,000. |
| `max_source_age_seconds` | Optional content-age limit with the same range. When present, every record needs a source timestamp. Omit the field to disable this check. |
| `records` | 1–1,000 objects with unique, nonblank `id` values of at most 128 characters and no surrounding whitespace. |
| `record.observed_at` | Declared time this exact source observation was captured. |
| `record.source_updated_at` | Declared content-update time represented in that observation. Requires an explicit source-age policy. |
| `record.source_id` | Nonblank source identifier, at most 512 characters. No source is fetched. |
| `record.evidence_id` | Nonblank reference to the observation, at most 512 characters. No evidence is authenticated. |

Accepted timestamp forms are `YYYY-MM-DDTHH:MM:SSZ` and a numeric timezone offset,
optionally with one to six fractional second digits. Date-only, timezone-free,
unknown `-00:00` offsets, invalid calendar dates, and leap seconds are rejected.
Offsets are normalized to UTC. Equality with an age limit passes; future times,
even by one microsecond, fail. Clock-skew tolerance is not implemented.

Missing or invalid row timestamps and evidence fields block that row. Source
updates later than observation time are inconsistent and also block it. Malformed
policies, unknown fields, duplicate IDs, or invalid batch sizes fail the run before
scoring. Supplying `source_updated_at` without a source-age policy is an input
error, even when the value is null. No rows are silently dropped.

Normalized input JSON is limited to 1,000,000 UTF-8 bytes. Raw documents and page
bodies are outside this contract. The platform parses JSON before the evaluator;
the Actor does not guarantee detection of duplicate JSON object keys.

### Output

The default dataset contains **one item for the whole batch**, with:

- `decision`: `pass` only when every record meets the policy.
- `summary`: total, passed, and blocked record counts plus diagnostic counts.
- `results`: record IDs, ages in seconds, recency statuses, evidence completeness,
  decisions, and diagnostic codes.
- `policy` and normalized `as_of`: the exact policy and evaluation clock used.
- `provenance_verified: false` and `truth_verified: false`.

The Summary view shows the batch counts. The Record diagnostics view contains the
nested `results` array and policy. Download JSON for the complete report.

**A completed evaluation succeeds as an Actor run even when its policy decision
is `fail`.** API and automation clients must inspect `decision` in the output;
successful run status alone does not mean records passed. Invalid batch input or
a storage failure fails the Actor run.

Recency status is `fresh`, `stale`, `unknown`, or `invalid`. A disabled source-age
check is `not_checked`. Missing or invalid times have null ages; future times may
have negative ages. Multiple diagnostics can apply to one record, so diagnostic
counts can exceed the number of blocked records.

### Using saved crawler metadata

For a saved [Website Content Crawler dataset](https://docs.apify.com/integrations/n8n/website-content-crawler),
prepare input with the following mapping:

| Saved field | Gate field |
| --- | --- |
| `crawl.loadedTime` | `observed_at` |
| `crawl.loadedUrl`, falling back to `url` | `source_id` |
| Your dataset ID plus the item's original offset | `evidence_id`, such as `dataset:ID:item:20` |
| An opaque unique row ID | `id` |

Preserve the original dataset order and offset when constructing evidence
references. These references remain declarations; the Actor does not open the
dataset to verify them. Supply `as_of` and `max_age_seconds` yourself.

Do not map `crawl.loadedTime` to `source_updated_at`. A load timestamp says when
the page was fetched. If reliable source-update metadata is unavailable, omit the
source-age policy or accept an explicit missing-source-time failure when that
policy is required. This Actor accepts prepared metadata; automatic dataset
downloading is not implemented.

### Data and cost

Apify receives and stores run input and output under its storage settings. Input
includes the source and evidence identifiers you supply. Output contains opaque
record IDs, policy metadata, and diagnostics; it does not echo source or evidence
identifiers. Use opaque IDs and avoid secrets in all input fields. No automatic
deletion or zero-retention behavior is promised.

There are no third-party model charges. Consult the current pricing shown before
running for platform execution/storage charges and any Actor fee. Local timing
is not evidence of cloud cost.

Memory is configured at 256 MB. The process has a 60-second deadline beginning
when the Python entrypoint starts, including SDK imports and storage operations.
Container startup is additional. Partial output is not guaranteed on timeout.
These limits bound individual work, not account spending across unlimited runs.

Version 0.1.1 adds Apify packaging and the saved-crawler metadata example to the
v0.1.0 deterministic evaluator. Provenance and truth verification remain absent.

# Actor input Schema

## `as_of` (type: `string`):

Explicit timezone-aware timestamp. The Actor never substitutes the current clock.

## `max_age_seconds` (type: `integer`):

Inclusive maximum age of observed\_at. Zero permits only an exact timestamp match.

## `max_source_age_seconds` (type: `integer`):

Optional. Omit to disable source-age checks. When supplied, source\_updated\_at is required on each record. Fetch time is not a content-update timestamp.

## `records` (type: `array`):

1–1,000 objects with unique nonblank IDs (up to 128 characters), observed\_at, source\_id, evidence\_id, and optional source\_updated\_at. Source/evidence identifiers are up to 512 characters. Missing or invalid row metadata returns diagnostics.

## Actor input object example

```json
{
  "as_of": "2026-09-24T12:00:00Z",
  "max_age_seconds": 86400,
  "max_source_age_seconds": 2592000,
  "records": [
    {
      "id": "healthy",
      "observed_at": "2026-09-24T11:00:00Z",
      "source_updated_at": "2026-09-23T12:00:00Z",
      "source_id": "feed-a",
      "evidence_id": "capture-1"
    },
    {
      "id": "old-observation",
      "observed_at": "2026-09-22T12:00:00Z",
      "source_updated_at": "2026-09-21T12:00:00Z",
      "source_id": "feed-a",
      "evidence_id": "capture-2"
    },
    {
      "id": "fresh-fetch-old-content",
      "observed_at": "2026-09-24T11:00:00Z",
      "source_updated_at": "2026-08-01T12:00:00Z",
      "source_id": "feed-a",
      "evidence_id": "capture-3"
    },
    {
      "id": "unknown-source-age",
      "observed_at": "2026-09-24T11:00:00Z",
      "source_updated_at": null,
      "source_id": "feed-a",
      "evidence_id": "capture-4"
    }
  ]
}
```

# Actor output Schema

## `report` (type: `string`):

One summary plus all record results. Provenance and truth are not verified.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "as_of": "2026-09-24T12:00:00Z",
    "max_age_seconds": 86400,
    "max_source_age_seconds": 2592000,
    "records": [
        {
            "id": "healthy",
            "observed_at": "2026-09-24T11:00:00Z",
            "source_updated_at": "2026-09-23T12:00:00Z",
            "source_id": "feed-a",
            "evidence_id": "capture-1"
        },
        {
            "id": "old-observation",
            "observed_at": "2026-09-22T12:00:00Z",
            "source_updated_at": "2026-09-21T12:00:00Z",
            "source_id": "feed-a",
            "evidence_id": "capture-2"
        },
        {
            "id": "fresh-fetch-old-content",
            "observed_at": "2026-09-24T11:00:00Z",
            "source_updated_at": "2026-08-01T12:00:00Z",
            "source_id": "feed-a",
            "evidence_id": "capture-3"
        },
        {
            "id": "unknown-source-age",
            "observed_at": "2026-09-24T11:00:00Z",
            "source_updated_at": null,
            "source_id": "feed-a",
            "evidence_id": "capture-4"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("saimislam/dataset-recency-evidence-gate").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "as_of": "2026-09-24T12:00:00Z",
    "max_age_seconds": 86400,
    "max_source_age_seconds": 2592000,
    "records": [
        {
            "id": "healthy",
            "observed_at": "2026-09-24T11:00:00Z",
            "source_updated_at": "2026-09-23T12:00:00Z",
            "source_id": "feed-a",
            "evidence_id": "capture-1",
        },
        {
            "id": "old-observation",
            "observed_at": "2026-09-22T12:00:00Z",
            "source_updated_at": "2026-09-21T12:00:00Z",
            "source_id": "feed-a",
            "evidence_id": "capture-2",
        },
        {
            "id": "fresh-fetch-old-content",
            "observed_at": "2026-09-24T11:00:00Z",
            "source_updated_at": "2026-08-01T12:00:00Z",
            "source_id": "feed-a",
            "evidence_id": "capture-3",
        },
        {
            "id": "unknown-source-age",
            "observed_at": "2026-09-24T11:00:00Z",
            "source_updated_at": None,
            "source_id": "feed-a",
            "evidence_id": "capture-4",
        },
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("saimislam/dataset-recency-evidence-gate").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "as_of": "2026-09-24T12:00:00Z",
  "max_age_seconds": 86400,
  "max_source_age_seconds": 2592000,
  "records": [
    {
      "id": "healthy",
      "observed_at": "2026-09-24T11:00:00Z",
      "source_updated_at": "2026-09-23T12:00:00Z",
      "source_id": "feed-a",
      "evidence_id": "capture-1"
    },
    {
      "id": "old-observation",
      "observed_at": "2026-09-22T12:00:00Z",
      "source_updated_at": "2026-09-21T12:00:00Z",
      "source_id": "feed-a",
      "evidence_id": "capture-2"
    },
    {
      "id": "fresh-fetch-old-content",
      "observed_at": "2026-09-24T11:00:00Z",
      "source_updated_at": "2026-08-01T12:00:00Z",
      "source_id": "feed-a",
      "evidence_id": "capture-3"
    },
    {
      "id": "unknown-source-age",
      "observed_at": "2026-09-24T11:00:00Z",
      "source_updated_at": null,
      "source_id": "feed-a",
      "evidence_id": "capture-4"
    }
  ]
}' |
apify call saimislam/dataset-recency-evidence-gate --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,saimislam/dataset-recency-evidence-gate"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/LIFHr81GdVj06rVs2/builds/aTsc1WSRnwTfRMigV/openapi.json
