Go to example tasks
Gate an AI Agent Release Against a Baseline
Created by
Dries Vd
Compare current AI-agent runs with a previous model or release. Detect reliability, quality, latency, cost-per-success, and margin regressions before production rollout.
AI Agent Performance Auditorwintry_nutmeg/ai-agent-performance-auditor
Release decision
Health score
Status
Performance metrics
+5 fieldsTextNumberBooleanListObject
Input
Current AI-agent runs(required)
id:run-101+2
success:true+2
costUsd:0.04+2
revenueUsd:0.2+2
latencyMs:1200+2
qualityScore:92+2
version:1.1+2
model:model-a+2
Optional baseline runs
id:run-001+1
success:true+1
costUsd:0.06+1
revenueUsd:0.2+1
latencyMs:1700+1
qualityScore:84+1
version:1.0+1
model:model-a+1
Optional regression thresholds
Output fields
Release decision
Health score
Status
Performance metrics
Regressions
Versions
Models
Recommended actions
Checked
Sign up on Apify01
Create your Apify account to access the AI Agent Performance Auditor.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
