# Skill Upgrade Evidence Check (`ysoserious_h/benchmark-evidence-preflight`) Actor

Review completed Skill evaluations for missing evidence and regressions. Each new user gets the first 5 delivered reports free once across runs, then $0.10/report, including blocked comparisons. JSON, Markdown and evidence ZIP. Free input exporter; no paid companion.

- **URL**: https://apify.com/ysoserious\_h/benchmark-evidence-preflight.md
- **Developed by:** [jw H](https://apify.com/ysoserious_h) (community)
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$100.00 / 1,000 evidence review reports

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Skill Upgrade Evidence Check

Review completed old/new or with/without Skill evaluation records before adopting a change. Check coverage against the original plan, locate reported regressions, and keep useful observations when an overall comparison is blocked.

### Bring an existing evaluation

You need the original `evals.json`, the completed iteration directory, explicit candidate and baseline directory names, and the number of repeats originally planned for each case and role. Two SKILL.md files or an aggregate score alone are not enough.

This checker does not execute your Skill, call a model, independently grade answers, verify a grader's claims, or make the release decision.

### Price and delivery choices

Each new user receives their **first five delivered reports free, once across all runs of this product**. After those five reports, the price is **$0.10 per delivered report**. The allowance belongs to the Apify account that starts the run; it does not reset per run or per month and is not an input-row allowance or Apify platform credit.

A delivered `BLOCKED_COMPARISON` report counts as one report, just like a complete comparison. Invalid input or unconfirmed delivery neither consumes a free report nor generates a report-fee event. A new run that delivers another report counts separately; retrying the same confirmed delivery must not consume another allowance or charge again. Free reports have no report-fee charge. Separate platform/account services are outside this report offer; review the platform's displayed charges before starting.

The input exporter is a **free, separate download**; no Capafy purchase is required to prepare Actor input. It collects and validates an existing evaluation snapshot but does not generate a local evaluation report.

Free input exporter: [Download version 1.0.0](https://github.com/Ysohjw/skill-evidence-input-exporter/releases/download/v1.0.0/skill-evidence-input-exporter-1.0.0.zip). The four input-only files are distributed under the MIT License; the hosted checker and report engine are not included.

If you already have compatible transport JSON, you can submit it directly. The optional Capafy subscription still requires private hosted acceptance and is not required to use this Actor.

### Prepare input locally

The free exporter uses Python 3.10+ and the standard library. Extract it to a short local path. From the `skill-evidence-input-exporter` directory, place a request file next to your original plan and iteration:

```json
{
  "iteration": "./iteration-2",
  "evals": "./evals.json",
  "candidate": "new_skill",
  "baseline": "old_skill",
  "runs": 2,
  "review_context": "Inspect reported regressions before adopting this candidate."
}
```

Use your actual original roles and repeat count. Relative paths are resolved from the request file. Export outside the iteration to a new file:

```text
python -X utf8 -B src/prepare_input.py --request "PATH/review-request.json" --out "PATH/new-input.json" --acknowledge-data
```

Read the exported JSON before uploading it. Exporting creates a local file and does not send it anywhere. Paste the **whole JSON object** into this Actor's JSON input, including `dataHandlingAcknowledged` and `evidence`. The form editor accepts those same two fields. A local path, URL or ZIP is not valid Actor input; this checker does not fetch evidence URLs or substitute sample data.

Default export includes the plan, recognized grading/evaluation/run/timing records, directory names and layout markers. It omits task artifact contents. Add `--include-text-artifacts` only if you intend to include ordinary JSON, TXT, MD or CSV files under `outputs/` or `artifacts/`. Binary files and executable scripts are not collected or executed.

### Empty startup

Starting without evidence, or with the unchanged empty default and data consent left unchecked, returns SETUP\_REQUIRED with one NOT\_EVALUATED preparation-guidance row in the default dataset. Open the empty-start preparation guidance in Output for missing inputs and next steps. It does not generate an evaluation report or ZIP, call report charging, or access the free-report allowance. Platform runtime costs may still apply. To request a review, provide completed evidence and explicitly acknowledge data handling. Nonempty invalid submissions are still rejected.

### Read the result

| Evidence status | Meaning |
| --- | --- |
| `READY_FOR_HUMAN_REVIEW` | The supplied records support a complete structural comparison. Inspect reported regressions and actual outputs. |
| `BLOCKED_COMPARISON` | Missing, extra or inconsistent evidence prevents the overall comparison. Valid observations remain available. This delivered report uses one free allowance, or is charged after the first five. |
| `SETUP_REQUIRED` | No evaluation was submitted. One NOT\_EVALUATED preparation-guidance row; no report, evidence ZIP, report-fee call or allowance use. |
| `INPUT_ERROR` | The input contract is not met. No comparison or evidence ZIP is produced, and no report-fee event is requested. |

The run's storage contains `OUTPUT` (JSON), `REPORT.md`, `REVIEW.zip` for valid evidence input, and `RUN_INFO` with completion status and file hashes. `BILLING_STATE` records the separate trial-allowance or paid report-fee outcome. It distinguishes a free delivery from a confirmed paid event; an event count alone is not a customer-payment receipt. An empty startup produces one preparation-guidance row in the default dataset, plus OUTPUT, RUN\_INFO and BILLING\_STATE with setup status; REPORT.md and REVIEW.zip are absent. The guidance row is not an evaluation or a paid report. If delivery stops early, some files can be absent. Platform success, comparison readiness and settled payment are different facts.

Download and extract `REVIEW.zip` before opening its relative links. The ZIP contains supplied evidence as well as reports. Omitted artifacts remain in your original workspace. Links locate files; they do not certify file contents. A higher reported mean can coexist with a regression. After a real correction and authorized rerun, the responsible grader must update the evidence before a new review can reflect the change.

#### Synthetic example

The included authored fixture has 2 cases, 2 repeats and 2 roles: 8/8 valid slots. Candidate mean is 87.5%, baseline mean 75%, and the difference is +12.5 percentage points, with one reported assertion regression. A separate missing-grade example retains seven valid observations and withholds the overall comparison. These are synthetic examples, not customer results, model evaluations or proof of improvement.

### Privacy and limits

The local CLI makes no network or model calls. Using this Actor sends the selected input to Apify and stores input, reports and included evidence there. Paths and text can contain private information or credentials. There is no automatic redaction: use an authorized, reviewed copy. If you ask an agent client to prepare or discuss files, its model and data policies also apply.

Download results you need and review your account's access, sharing, retention and deletion controls. Shared links may expose data to their holders. Trial accounting uses a persistent product-specific record keyed to the platform user identity, with report delivery identifiers and allowance use; it does not need to copy evidence into that record. Evidence retention is separate from trial accounting. The checker does not delete cloud records automatically and does not promise a fixed retention period for every account. It does not fetch external evidence or send it to an LLM.

Limits: 8 MiB serialized JSON, 1,000 selected files, 3,000 directories, 2 MiB per file, and 10,000 inspected directory entries. The core supports up to 200 cases and 1–20 planned repeats. Unsafe, linked/reparse, ambiguous and case-colliding paths are rejected; files are not silently truncated. Timing and token fields describe supplied measurements; output characters are not converted to tokens or costs.

Independent tool; references to Apify, Anthropic or OpenAI do not imply endorsement. The maintainer remains responsible for the upgrade decision.

### Support

Contact: hjwysoserious@gmail.com. Include the run ID and a sanitized description; do not email tokens or private evaluation files. No fixed response-time guarantee is offered.

# Actor input Schema

## `dataHandlingAcknowledged` (type: `boolean`):

Must be true. Exporting locally does not authorize uploading. Prompts, evidence text, metadata and paths can contain sensitive data; automatic redaction is not provided.

## `evidence` (type: `object`):

Paste the evidence object from your export. Requires original evals.json text, explicit candidate/baseline/repeats and ordinary relative directory/file lists. Max 8 MiB JSON, 1000 files, 3000 directories and 2 MiB per file. Artifact contents are omitted by default. Leave this empty only for setup guidance; a review request requires completed evidence and explicit data consent. Empty objects may be omitted by the native form.

## Actor input object example

```json
{}
```

# Actor output Schema

## `files` (type: `string`):

No description

## `setup_guidance` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("ysoserious_h/benchmark-evidence-preflight").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("ysoserious_h/benchmark-evidence-preflight").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call ysoserious_h/benchmark-evidence-preflight --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ysoserious_h/benchmark-evidence-preflight"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/RZO0EUjfI3xR006cs/builds/tl6imS0Sxx8WhblpX/openapi.json
