# rag-regression-evaluator (`saimislam/rag-regression-evaluator`) Actor

Compare baseline and candidate RAG outputs using lexical metrics. Get per-case score changes, regression flags, and a pass/fail threshold decision for up to 100 test cases. No LLM calls. Experimental: does not verify factual accuracy or semantic correctness.

- **URL**: https://apify.com/saimislam/rag-regression-evaluator.md
- **Developed by:** [Saim Islam](https://apify.com/saimislam) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## RAG Change Diagnostics

Compare saved baseline and candidate RAG traces using deterministic lexical
metrics. Get per-case score changes and an aggregate threshold decision.
No LLM or embedding API is called.

**Experimental lexical diagnostic. A pass does not establish factual accuracy or
semantic correctness.** Paraphrases, negation, numbers, and entity roles can be
misclassified. Do not use this as the sole production release check.

### Example

```json
{
  "cases": [{
    "id": "capital-france",
    "query": "What is the capital of France?",
    "reference_answer": "Paris",
    "baseline": {
      "contexts": ["Paris is the capital of France."],
      "answer": "The capital of France is Paris."
    },
    "candidate": {
      "contexts": ["Paris is the capital of France."],
      "answer": "The capital of France is Berlin."
    }
  }],
  "regression_threshold": 0.15
}
```

This example produces `release_decision: "fail"` and a case delta of `-0.4167`.
An unchanged correct answer produces a delta of zero and passes. Use the input
form's example for a two-case demonstration.

### Input limits

- 1–100 cases, with unique nonempty IDs of at most 128 characters.
- Each case needs a query, baseline, and candidate. Reference answer is optional.
- Each output requires `contexts` (0–20 strings) and `answer` (a string).
- Query: at most 4,000 characters. Each answer, reference, or context: 20,000.
- Normalized JSON: at most 1,000,000 UTF-8 bytes, excluding optional whitespace.
- Threshold: a finite JSON number from 0.000001 to 1; default 0.15.
- Empty answers and empty context arrays are accepted to diagnose missing output.
- Unknown fields, duplicate IDs, or malformed values fail the run before scoring.

The program exits after 60 seconds from application startup if unfinished. Cloud
startup and Python import time are additional. Partial output is not guaranteed on timeout. Memory
is configured at 256 MB. These limits bound individual work, not total account
spending across unlimited runs.

### Reading results

The default dataset contains one summary with `results`, `cases_evaluated`,
`regressions_detected`, `average_score_delta`, and `release_decision`.
`evaluation_mode` is `lexical_diagnostic` and `semantic_correctness_verified` is
always false. Metrics named grounding and reference coverage mean token overlap.

The Actor succeeds when evaluation finishes, even if the evaluated release
decision is `fail`. API/CI clients must inspect `release_decision`; successful
HTTP or Actor status alone does not mean the candidate passed. Any case whose
score drops by at least the threshold fails the aggregate decision. Individual
failure codes alone do not necessarily fail it.

The score weights query coverage, answer/context overlap, and optional reference
coverage, with a small repeated-context penalty. It uses English stopwords and
discards one-character tokens; numerical and multilingual interpretation is weak.

### Data and cost

Apify receives and stores run input and output according to its storage settings.
Do not submit secrets or sensitive documents. Output retains supplied case IDs
and numeric diagnostics, not raw queries, contexts, or answers. Use opaque IDs.
The code sends no data to an external model or analytics service. This is not a
zero-retention service, and no automatic deletion schedule is promised.

There are no third-party model charges. Platform execution/storage charges and
any Actor fee depend on the current Apify pricing shown before running. A rounded
`$0.000` result from a tiny test is not a guarantee that runs are free.

Version 0.2.1 adds validation and process limits; scoring remains the v0.2 lexical
baseline. Zero thresholds are now rejected to avoid treating unchanged scores
as regressions. Semantic limitations remain unresolved.

# Actor input Schema

## `cases` (type: `array`):

1–100 cases. Unique ID (128 chars), query (4000 chars), baseline/candidate with answer and 0–20 contexts. Each answer/reference/context at most 20000 chars. See README for full contract.

## `regression_threshold` (type: `number`):

Score-drop threshold from 0.000001 to 1. Zero is rejected.

## Actor input object example

```json
{
  "cases": [
    {
      "id": "capital-france",
      "query": "What is the capital of France?",
      "reference_answer": "Paris",
      "baseline": {
        "contexts": [
          "France is a country in Western Europe. Its capital is Paris."
        ],
        "answer": "The capital of France is Paris."
      },
      "candidate": {
        "contexts": [
          "France is a country in Western Europe. Its capital is Paris."
        ],
        "answer": "The capital of France is Berlin."
      }
    },
    {
      "id": "largest-planet",
      "query": "What is the largest planet in the Solar System?",
      "reference_answer": "Jupiter",
      "baseline": {
        "contexts": [
          "Jupiter is the largest planet in the Solar System."
        ],
        "answer": "Jupiter is the largest planet in the Solar System."
      },
      "candidate": {
        "contexts": [
          "Jupiter is the largest planet in the Solar System."
        ],
        "answer": "Jupiter is the largest planet in the Solar System."
      }
    }
  ],
  "regression_threshold": 0.15
}
```

# Actor output Schema

## `results` (type: `string`):

Default dataset: one summary with counts, a lexical threshold decision, and per-case metrics. Inspect release\_decision; successful Actor execution alone does not mean the candidate passed.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "cases": [
        {
            "id": "capital-france",
            "query": "What is the capital of France?",
            "reference_answer": "Paris",
            "baseline": {
                "contexts": [
                    "France is a country in Western Europe. Its capital is Paris."
                ],
                "answer": "The capital of France is Paris."
            },
            "candidate": {
                "contexts": [
                    "France is a country in Western Europe. Its capital is Paris."
                ],
                "answer": "The capital of France is Berlin."
            }
        },
        {
            "id": "largest-planet",
            "query": "What is the largest planet in the Solar System?",
            "reference_answer": "Jupiter",
            "baseline": {
                "contexts": [
                    "Jupiter is the largest planet in the Solar System."
                ],
                "answer": "Jupiter is the largest planet in the Solar System."
            },
            "candidate": {
                "contexts": [
                    "Jupiter is the largest planet in the Solar System."
                ],
                "answer": "Jupiter is the largest planet in the Solar System."
            }
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("saimislam/rag-regression-evaluator").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "cases": [
        {
            "id": "capital-france",
            "query": "What is the capital of France?",
            "reference_answer": "Paris",
            "baseline": {
                "contexts": ["France is a country in Western Europe. Its capital is Paris."],
                "answer": "The capital of France is Paris.",
            },
            "candidate": {
                "contexts": ["France is a country in Western Europe. Its capital is Paris."],
                "answer": "The capital of France is Berlin.",
            },
        },
        {
            "id": "largest-planet",
            "query": "What is the largest planet in the Solar System?",
            "reference_answer": "Jupiter",
            "baseline": {
                "contexts": ["Jupiter is the largest planet in the Solar System."],
                "answer": "Jupiter is the largest planet in the Solar System.",
            },
            "candidate": {
                "contexts": ["Jupiter is the largest planet in the Solar System."],
                "answer": "Jupiter is the largest planet in the Solar System.",
            },
        },
    ] }

# Run the Actor and wait for it to finish
run = client.actor("saimislam/rag-regression-evaluator").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "cases": [
    {
      "id": "capital-france",
      "query": "What is the capital of France?",
      "reference_answer": "Paris",
      "baseline": {
        "contexts": [
          "France is a country in Western Europe. Its capital is Paris."
        ],
        "answer": "The capital of France is Paris."
      },
      "candidate": {
        "contexts": [
          "France is a country in Western Europe. Its capital is Paris."
        ],
        "answer": "The capital of France is Berlin."
      }
    },
    {
      "id": "largest-planet",
      "query": "What is the largest planet in the Solar System?",
      "reference_answer": "Jupiter",
      "baseline": {
        "contexts": [
          "Jupiter is the largest planet in the Solar System."
        ],
        "answer": "Jupiter is the largest planet in the Solar System."
      },
      "candidate": {
        "contexts": [
          "Jupiter is the largest planet in the Solar System."
        ],
        "answer": "Jupiter is the largest planet in the Solar System."
      }
    }
  ]
}' |
apify call saimislam/rag-regression-evaluator --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,saimislam/rag-regression-evaluator"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/42b0d67uSxnyyfbVQ/builds/dZCaqiuLwuq3fyQw7/openapi.json
