rag-regression-evaluator avatar

rag-regression-evaluator

Pricing

Pay per usage

Go to Apify Store
rag-regression-evaluator

rag-regression-evaluator

Compare baseline and candidate RAG outputs using lexical metrics. Get per-case score changes, regression flags, and a pass/fail threshold decision for up to 100 test cases. No LLM calls. Experimental: does not verify factual accuracy or semantic correctness.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Saim Islam

Saim Islam

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 days ago

Last modified

Categories

Share

RAG Change Diagnostics

Compare saved baseline and candidate RAG traces using deterministic lexical metrics. Get per-case score changes and an aggregate threshold decision. No LLM or embedding API is called.

Experimental lexical diagnostic. A pass does not establish factual accuracy or semantic correctness. Paraphrases, negation, numbers, and entity roles can be misclassified. Do not use this as the sole production release check.

Example

{
"cases": [{
"id": "capital-france",
"query": "What is the capital of France?",
"reference_answer": "Paris",
"baseline": {
"contexts": ["Paris is the capital of France."],
"answer": "The capital of France is Paris."
},
"candidate": {
"contexts": ["Paris is the capital of France."],
"answer": "The capital of France is Berlin."
}
}],
"regression_threshold": 0.15
}

This example produces release_decision: "fail" and a case delta of -0.4167. An unchanged correct answer produces a delta of zero and passes. Use the input form's example for a two-case demonstration.

Input limits

  • 1–100 cases, with unique nonempty IDs of at most 128 characters.
  • Each case needs a query, baseline, and candidate. Reference answer is optional.
  • Each output requires contexts (0–20 strings) and answer (a string).
  • Query: at most 4,000 characters. Each answer, reference, or context: 20,000.
  • Normalized JSON: at most 1,000,000 UTF-8 bytes, excluding optional whitespace.
  • Threshold: a finite JSON number from 0.000001 to 1; default 0.15.
  • Empty answers and empty context arrays are accepted to diagnose missing output.
  • Unknown fields, duplicate IDs, or malformed values fail the run before scoring.

The program exits after 60 seconds from application startup if unfinished. Cloud startup and Python import time are additional. Partial output is not guaranteed on timeout. Memory is configured at 256 MB. These limits bound individual work, not total account spending across unlimited runs.

Reading results

The default dataset contains one summary with results, cases_evaluated, regressions_detected, average_score_delta, and release_decision. evaluation_mode is lexical_diagnostic and semantic_correctness_verified is always false. Metrics named grounding and reference coverage mean token overlap.

The Actor succeeds when evaluation finishes, even if the evaluated release decision is fail. API/CI clients must inspect release_decision; successful HTTP or Actor status alone does not mean the candidate passed. Any case whose score drops by at least the threshold fails the aggregate decision. Individual failure codes alone do not necessarily fail it.

The score weights query coverage, answer/context overlap, and optional reference coverage, with a small repeated-context penalty. It uses English stopwords and discards one-character tokens; numerical and multilingual interpretation is weak.

Data and cost

Apify receives and stores run input and output according to its storage settings. Do not submit secrets or sensitive documents. Output retains supplied case IDs and numeric diagnostics, not raw queries, contexts, or answers. Use opaque IDs. The code sends no data to an external model or analytics service. This is not a zero-retention service, and no automatic deletion schedule is promised.

There are no third-party model charges. Platform execution/storage charges and any Actor fee depend on the current Apify pricing shown before running. A rounded $0.000 result from a tiny test is not a guarantee that runs are free.

Version 0.2.1 adds validation and process limits; scoring remains the v0.2 lexical baseline. Zero thresholds are now rejected to avoid treating unchanged scores as regressions. Semantic limitations remain unresolved.