Skill Upgrade Evidence Check avatar

Skill Upgrade Evidence Check

Under maintenance

Pricing

$100.00 / 1,000 evidence review reports

Go to Apify Store
Skill Upgrade Evidence Check

Skill Upgrade Evidence Check

Under maintenance

Review completed Skill evaluations for missing evidence and regressions. Each new user gets the first 5 delivered reports free once across runs, then $0.10/report, including blocked comparisons. JSON, Markdown and evidence ZIP. Free input exporter; no paid companion.

Pricing

$100.00 / 1,000 evidence review reports

Rating

0.0

(0)

Developer

jw H

jw H

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

4 days ago

Last modified

Categories

Share

Review completed old/new or with/without Skill evaluation records before adopting a change. Check coverage against the original plan, locate reported regressions, and keep useful observations when an overall comparison is blocked.

Bring an existing evaluation

You need the original evals.json, the completed iteration directory, explicit candidate and baseline directory names, and the number of repeats originally planned for each case and role. Two SKILL.md files or an aggregate score alone are not enough.

This checker does not execute your Skill, call a model, independently grade answers, verify a grader's claims, or make the release decision.

Price and delivery choices

Each new user receives their first five delivered reports free, once across all runs of this product. After those five reports, the price is $0.10 per delivered report. The allowance belongs to the Apify account that starts the run; it does not reset per run or per month and is not an input-row allowance or Apify platform credit.

A delivered BLOCKED_COMPARISON report counts as one report, just like a complete comparison. Invalid input or unconfirmed delivery neither consumes a free report nor generates a report-fee event. A new run that delivers another report counts separately; retrying the same confirmed delivery must not consume another allowance or charge again. Free reports have no report-fee charge. Separate platform/account services are outside this report offer; review the platform's displayed charges before starting.

The input exporter is a free, separate download; no Capafy purchase is required to prepare Actor input. It collects and validates an existing evaluation snapshot but does not generate a local evaluation report.

Free input exporter: Download version 1.0.0. The four input-only files are distributed under the MIT License; the hosted checker and report engine are not included.

If you already have compatible transport JSON, you can submit it directly. The optional Capafy subscription still requires private hosted acceptance and is not required to use this Actor.

Prepare input locally

The free exporter uses Python 3.10+ and the standard library. Extract it to a short local path. From the skill-evidence-input-exporter directory, place a request file next to your original plan and iteration:

{
"iteration": "./iteration-2",
"evals": "./evals.json",
"candidate": "new_skill",
"baseline": "old_skill",
"runs": 2,
"review_context": "Inspect reported regressions before adopting this candidate."
}

Use your actual original roles and repeat count. Relative paths are resolved from the request file. Export outside the iteration to a new file:

python -X utf8 -B src/prepare_input.py --request "PATH/review-request.json" --out "PATH/new-input.json" --acknowledge-data

Read the exported JSON before uploading it. Exporting creates a local file and does not send it anywhere. Paste the whole JSON object into this Actor's JSON input, including dataHandlingAcknowledged and evidence. The form editor accepts those same two fields. A local path, URL or ZIP is not valid Actor input; this checker does not fetch evidence URLs or substitute sample data.

Default export includes the plan, recognized grading/evaluation/run/timing records, directory names and layout markers. It omits task artifact contents. Add --include-text-artifacts only if you intend to include ordinary JSON, TXT, MD or CSV files under outputs/ or artifacts/. Binary files and executable scripts are not collected or executed.

Empty startup

Starting without evidence, or with the unchanged empty default and data consent left unchecked, returns SETUP_REQUIRED with one NOT_EVALUATED preparation-guidance row in the default dataset. Open the empty-start preparation guidance in Output for missing inputs and next steps. It does not generate an evaluation report or ZIP, call report charging, or access the free-report allowance. Platform runtime costs may still apply. To request a review, provide completed evidence and explicitly acknowledge data handling. Nonempty invalid submissions are still rejected.

Read the result

Evidence statusMeaning
READY_FOR_HUMAN_REVIEWThe supplied records support a complete structural comparison. Inspect reported regressions and actual outputs.
BLOCKED_COMPARISONMissing, extra or inconsistent evidence prevents the overall comparison. Valid observations remain available. This delivered report uses one free allowance, or is charged after the first five.
SETUP_REQUIREDNo evaluation was submitted. One NOT_EVALUATED preparation-guidance row; no report, evidence ZIP, report-fee call or allowance use.
INPUT_ERRORThe input contract is not met. No comparison or evidence ZIP is produced, and no report-fee event is requested.

The run's storage contains OUTPUT (JSON), REPORT.md, REVIEW.zip for valid evidence input, and RUN_INFO with completion status and file hashes. BILLING_STATE records the separate trial-allowance or paid report-fee outcome. It distinguishes a free delivery from a confirmed paid event; an event count alone is not a customer-payment receipt. An empty startup produces one preparation-guidance row in the default dataset, plus OUTPUT, RUN_INFO and BILLING_STATE with setup status; REPORT.md and REVIEW.zip are absent. The guidance row is not an evaluation or a paid report. If delivery stops early, some files can be absent. Platform success, comparison readiness and settled payment are different facts.

Download and extract REVIEW.zip before opening its relative links. The ZIP contains supplied evidence as well as reports. Omitted artifacts remain in your original workspace. Links locate files; they do not certify file contents. A higher reported mean can coexist with a regression. After a real correction and authorized rerun, the responsible grader must update the evidence before a new review can reflect the change.

Synthetic example

The included authored fixture has 2 cases, 2 repeats and 2 roles: 8/8 valid slots. Candidate mean is 87.5%, baseline mean 75%, and the difference is +12.5 percentage points, with one reported assertion regression. A separate missing-grade example retains seven valid observations and withholds the overall comparison. These are synthetic examples, not customer results, model evaluations or proof of improvement.

Privacy and limits

The local CLI makes no network or model calls. Using this Actor sends the selected input to Apify and stores input, reports and included evidence there. Paths and text can contain private information or credentials. There is no automatic redaction: use an authorized, reviewed copy. If you ask an agent client to prepare or discuss files, its model and data policies also apply.

Download results you need and review your account's access, sharing, retention and deletion controls. Shared links may expose data to their holders. Trial accounting uses a persistent product-specific record keyed to the platform user identity, with report delivery identifiers and allowance use; it does not need to copy evidence into that record. Evidence retention is separate from trial accounting. The checker does not delete cloud records automatically and does not promise a fixed retention period for every account. It does not fetch external evidence or send it to an LLM.

Limits: 8 MiB serialized JSON, 1,000 selected files, 3,000 directories, 2 MiB per file, and 10,000 inspected directory entries. The core supports up to 200 cases and 1–20 planned repeats. Unsafe, linked/reparse, ambiguous and case-colliding paths are rejected; files are not silently truncated. Timing and token fields describe supplied measurements; output characters are not converted to tokens or costs.

Independent tool; references to Apify, Anthropic or OpenAI do not imply endorsement. The maintainer remains responsible for the upgrade decision.

Support

Contact: hjwysoserious@gmail.com. Include the run ID and a sanitized description; do not email tokens or private evaluation files. No fixed response-time guarantee is offered.