rag-regression-evaluator
Pricing
Pay per usage
rag-regression-evaluator
Compare baseline and candidate RAG outputs using lexical metrics. Get per-case score changes, regression flags, and a pass/fail threshold decision for up to 100 test cases. No LLM calls. Experimental: does not verify factual accuracy or semantic correctness.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Saim Islam
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
6 days ago
Last modified
Categories
Share
RAG Change Diagnostics
Compare saved baseline and candidate RAG traces using deterministic lexical metrics. Get per-case score changes and an aggregate threshold decision. No LLM or embedding API is called.
Experimental lexical diagnostic. A pass does not establish factual accuracy or semantic correctness. Paraphrases, negation, numbers, and entity roles can be misclassified. Do not use this as the sole production release check.
Example
{"cases": [{"id": "capital-france","query": "What is the capital of France?","reference_answer": "Paris","baseline": {"contexts": ["Paris is the capital of France."],"answer": "The capital of France is Paris."},"candidate": {"contexts": ["Paris is the capital of France."],"answer": "The capital of France is Berlin."}}],"regression_threshold": 0.15}
This example produces release_decision: "fail" and a case delta of -0.4167.
An unchanged correct answer produces a delta of zero and passes. Use the input
form's example for a two-case demonstration.
Input limits
- 1–100 cases, with unique nonempty IDs of at most 128 characters.
- Each case needs a query, baseline, and candidate. Reference answer is optional.
- Each output requires
contexts(0–20 strings) andanswer(a string). - Query: at most 4,000 characters. Each answer, reference, or context: 20,000.
- Normalized JSON: at most 1,000,000 UTF-8 bytes, excluding optional whitespace.
- Threshold: a finite JSON number from 0.000001 to 1; default 0.15.
- Empty answers and empty context arrays are accepted to diagnose missing output.
- Unknown fields, duplicate IDs, or malformed values fail the run before scoring.
The program exits after 60 seconds from application startup if unfinished. Cloud startup and Python import time are additional. Partial output is not guaranteed on timeout. Memory is configured at 256 MB. These limits bound individual work, not total account spending across unlimited runs.
Reading results
The default dataset contains one summary with results, cases_evaluated,
regressions_detected, average_score_delta, and release_decision.
evaluation_mode is lexical_diagnostic and semantic_correctness_verified is
always false. Metrics named grounding and reference coverage mean token overlap.
The Actor succeeds when evaluation finishes, even if the evaluated release
decision is fail. API/CI clients must inspect release_decision; successful
HTTP or Actor status alone does not mean the candidate passed. Any case whose
score drops by at least the threshold fails the aggregate decision. Individual
failure codes alone do not necessarily fail it.
The score weights query coverage, answer/context overlap, and optional reference coverage, with a small repeated-context penalty. It uses English stopwords and discards one-character tokens; numerical and multilingual interpretation is weak.
Data and cost
Apify receives and stores run input and output according to its storage settings. Do not submit secrets or sensitive documents. Output retains supplied case IDs and numeric diagnostics, not raw queries, contexts, or answers. Use opaque IDs. The code sends no data to an external model or analytics service. This is not a zero-retention service, and no automatic deletion schedule is promised.
There are no third-party model charges. Platform execution/storage charges and
any Actor fee depend on the current Apify pricing shown before running. A rounded
$0.000 result from a tiny test is not a guarantee that runs are free.
Version 0.2.1 adds validation and process limits; scoring remains the v0.2 lexical baseline. Zero thresholds are now rejected to avoid treating unchanged scores as regressions. Semantic limitations remain unresolved.